【问题标题】:spark scala typesafe config safe iterate over value of a specific column namespark scala typesafe config安全迭代特定列名的值
【发布时间】:2017-07-20 21:51:57
【问题描述】:

我在 Stackoverflow 上找到了类似的帖子。但是,我无法解决我的问题所以,这就是我写这篇文章的原因。

瞄准

目的是在加载 SQL 表(我使用 SQL Server)时执行列投影 [projection = filter columns]。

根据 scala 食谱,这是 [使用数组] 过滤列的方法:

sqlContext.read.jdbc(url,"person",Array("gender='M'"),prop)

但是,我不想在我的 Scala 代码中硬编码 Array("col1", "col2", ...) 这就是为什么我使用具有类型安全的配置文件(见下文)。

配置文件

dataset {
    type = sql
    sql{
        url = "jdbc://host:port:user:name:password"
        tablename = "ClientShampooBusinesLimited"
        driver = "driver"
        other = "i have a lot of other single string elements in the config file..."
        columnList = [
        {
            colname = "id"
            colAlias = "identifient"
        }
        {
            colname = "name"
            colAlias = "nom client"
        }
        {
            colname = "age"
            colAlias = "âge client"
        }
        ]
    }
}

让我们关注“columnList”:SQL 列的名称与“colname”完全对应。 'colAlias' 是我稍后将使用的字段。

data.scala 文件

lazy val columnList = configFromFile.getList("dataset.sql.columnList")
lazy val dbUrl = configFromFile.getList("dataset.sql.url")
lazy val DbTableName= configFromFile.getList("dataset.sql.tablename")
lazy val DriverName= configFromFile.getList("dataset.sql.driver")

configFromFile 是我自己在另一个自定义类中创建的。但这没关系。 columnList 的类型是“ConfigList”,这个类型来自于 typesafe。

主文件

def loadDataSQL(): DataFrame = {

val url = datasetConfig.dbUrl 
val dbTablename = datasetConfig.DbTableName
val dbDriver = datasetConfig.DriverName
val columns = // I need help to solve this


/* EDIT 2 march 2017
   This code should not be used. Have a look at the accepted answer.
*/
sparkSession.read.format("jdbc").options(
    Map("url" -> url,
    "dbtable" -> dbTablename,
    "predicates" -> columns,
    "driver" -> dbDriver))
    .load()
}

所以我所有的问题都是提取“colnames”值,以便将它们放入合适的数组中。有人可以帮我写出“val columns”的正确操作吗?

谢谢

【问题讨论】:

    标签: arrays scala apache-spark typesafe


    【解决方案1】:

    如果您正在寻找一种将 colname 值列表读入 Scala 数组的方法 - 我认为可以这样做:

    import scala.collection.JavaConverters._
    
    val columnList = configFromFile.getConfigList("dataset.sql.columnList")
    val colNames: Array[String] = columnList.asScala.map(_.getString("colname")).toArray
    

    使用提供的文件,这将导致Array(id, name, age)

    编辑: 至于您的实际目标,我实际上不知道任何名为 predication 的选项(我也无法在源代码中找到证据,使用 Spark 2.0.2)。

    JDBC 数据源根据使用的查询中选择的实际列执行“投影下推”。换句话说 - 只有 selected 列将从 DB 中读取,因此您可以在创建 DF 后立即在 select 中使用 colNames 数组,例如:

    import org.apache.spark.sql.functions._
    
    sparkSession.read
      .format("jdbc")
      .options(Map("url" -> url, "dbtable" -> dbTablename, "driver" -> dbDriver))
      .load()
      .select(colNames.map(col): _*) // selecting only desired columns
    

    【讨论】:

    • 亲爱的 Tzach Zohar,这正是我想要的。非常感谢您的帮助。
    • 但是,我在“predication”-> 列中出现错误,它显示“overloaded”方法。你知道是什么问题吗?谢谢
    • 不确定您所指的错误,但我更新了我的答案,希望能帮助您实现仅从 DB 中读取选定列的实际目标
    • 您好,很抱歉造成混淆,它不是“谓词”而是“谓词”。我会编辑帖子。但是,您提供给我的解决方案非常好。它就像一个魅力。非常感谢您提供有关“投影下推”的信息。乍一看我很不情愿,因为我认为它会先加载所有列,然后再进行投影。但是,现在我有了“投影下推”的概念。问候。
    猜你喜欢
    • 2014-01-02
    • 2018-03-12
    • 2021-07-01
    • 2019-07-02
    • 1970-01-01
    • 2019-01-06
    • 2015-01-14
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多