【问题标题】:How to do left outer join in spark sql?如何在spark sql中进行左外连接?
【发布时间】:2017-03-18 05:02:03
【问题描述】:

我正在尝试在 spark (1.6.2) 中进行左外连接,但它不起作用。我的sql查询是这样的:

sqlContext.sql("select t.type, t.uuid, p.uuid
from symptom_type t LEFT JOIN plugin p 
ON t.uuid = p.uuid 
where t.created_year = 2016 
and p.created_year = 2016").show()

结果是这样的:

+--------------------+--------------------+--------------------+
|                type|                uuid|                uuid|
+--------------------+--------------------+--------------------+
|              tained|89759dcc-50c0-490...|89759dcc-50c0-490...|
|             swapper|740cd0d4-53ee-438...|740cd0d4-53ee-438...|

我使用 LEFT JOIN 或 LEFT OUTER JOIN 得到了相同的结果(第二个 uuid 不为空)。

我希望第二个 uuid 列仅为 null。如何正确进行左外连接?

=== 附加信息 ==

如果我使用数据框进行左外连接,我得到了正确的结果。

s = sqlCtx.sql('select * from symptom_type where created_year = 2016')
p = sqlCtx.sql('select * from plugin where created_year = 2016')

s.join(p, s.uuid == p.uuid, 'left_outer')
.select(s.type, s.uuid.alias('s_uuid'), 
        p.uuid.alias('p_uuid'), s.created_date, p.created_year, p.created_month).show()

我得到这样的结果:

+-------------------+--------------------+-----------------+--------------------+------------+-------------+
|               type|              s_uuid|           p_uuid|        created_date|created_year|created_month|
+-------------------+--------------------+-----------------+--------------------+------------+-------------+
|             tained|6d688688-96a4-341...|             null|2016-01-28 00:27:...|        null|         null|
|             tained|6d688688-96a4-341...|             null|2016-01-28 00:27:...|        null|         null|
|             tained|6d688688-96a4-341...|             null|2016-01-28 00:27:...|        null|         null|

谢谢,

【问题讨论】:

    标签: apache-spark pyspark apache-spark-sql


    【解决方案1】:

    我认为您只需要使用 LEFT OUTER JOIN 而不是 LEFT JOIN 关键字即可。更多信息请查看Spark documentation

    【讨论】:

    • 我试过了,它仍然显示类似于内部连接的结果。第二个 uuid 上没有 null。
    【解决方案2】:

    我在您的代码中没有发现任何问题。 “左连接”或“左外连接”都可以正常工作。请再次检查数据,您显示的数据是匹配的。

    您还可以使用以下方法执行 Spark SQL 连接:

    // 左外连接显式

    df1.join(df2, df1["col1"] == df2["col1"], "left_outer")
    

    【讨论】:

    • 我添加了使用数据框的结果。我认为第二个 uuid 不是来自 sql 查询的 null 对我来说是个问题。看起来它只是执行内部连接
    • 请检查以下内容:-
    • 应该有两个==,而不是三个
    • 语法因语言而异,对于 pySpark '==' 是必需的,对于 Scala '===' (三个 = )是必需的。示例 - df1.join(df2, $"df1Key" === $"df2Key")
    • 他的代码有一个明显的问题。他在连接的右表中的列的 where 子句中包含过滤条件,从而产生内部连接。这是初学者常见的 SQL 错误。请参阅stackoverflow.com/questions/3256304/… 了解更多信息。
    【解决方案3】:

    您正在过滤掉p.created_year(和p.uuid)的空值

    where t.created_year = 2016 
    and p.created_year = 2016
    

    避免这种情况的方法是将p 的过滤子句移至ON 语句:

    sqlContext.sql("select t.type, t.uuid, p.uuid
    from symptom_type t LEFT JOIN plugin p 
    ON t.uuid = p.uuid 
    and p.created_year = 2016
    where t.created_year = 2016").show()
    

    这是正确的,但效率低下,因为我们还需要在连接发生之前过滤t.created_year。所以建议使用子查询:

    sqlContext.sql("select t.type, t.uuid, p.uuid
    from (
      SELECT type, uuid FROM symptom_type WHERE created_year = 2016 
    ) t LEFT JOIN (
      SELECT uuid FROM plugin WHERE created_year = 2016
    ) p 
    ON t.uuid = p.uuid").show()    
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-10-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2015-04-13
      相关资源
      最近更新 更多