【问题标题】:joining DataFrames in spark在 Spark 中加入 DataFrame
【发布时间】:2016-10-18 22:17:54
【问题描述】:

我想使用or 函数通过两个键连接两个数据框:edges 和 selectedComponent

 val selectedComponent = hiveContext.sql(s"""select * from $tableWithComponents
         |where component=$component""".stripMargin)

但不是这样

val theSelectedComponentEdges = hiveContext.sql(
  s"""select * from $tableWithComponents a join $edges b where (b.src=a.id or b.dst=a.id)""")

但使用连接功能

edges.join(selectedComponent, edges("src")===selectedComponent("id"))

但我不确定我应该如何在这里使用“或”。

任何人都可以帮助我:-)?

【问题讨论】:

    标签: scala apache-spark spark-dataframe


    【解决方案1】:
    edges.join(selectedComponent, (edges("src")===selectedComponent("id")) ||  (edges("dst")===selectedComponent("id")))
    

    【讨论】:

    • 正确,但列名有点混淆 - 应该是 (edges("src")===selectedComponent("id")) || (edges("dst")===selectedComponent("id"))
    • 我检查了val a = edges.join(selectedComponent, edges("src")===selectedComponent("id")) // val b = edges.join(selectedComponent, edges("dst")===selectedComponent("id")) // val theSelectedComponentEdges = a.unionAll(b),它更快,是一样的吗?
    • 结果是一样的,但是催化剂确实会用复杂的连接表达式搞乱连接(它经常做一个笛卡尔积......)。但我认为这是一个催化剂问题,您的优化虽然目前有效,但(希望)不是面向未来的。编辑:您实际上应该检查src == dst,您的实现中会有重复,但(我认为)不是第一个。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2016-08-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-08-27
    • 2021-06-11
    相关资源
    最近更新 更多