【问题标题】:How to filter data from a dataframe using pyspark如何使用 pyspark 从数据框中过滤数据
【发布时间】:2018-02-13 14:52:16
【问题描述】:

我有一个名为 mytable 的表作为可用的数据框,下面是该表

[+---+----+----+----+ | x|是| z| w| +---+----+----+----+ | 1|一个|空|空| | 1|空| b|空| | 1|空|空| c| | 2| d|空|空| | 2|空| e|空| | 2|空|空| f| +---+----+----+----+]

我想要我们按 col x 分组并连接 col y,z,w 的结果的结果。结果如下所示。

[+---+----+----+- | x|结果| +---+----+----+ | 1| a b c | | 2| d e f | +---+----+---+|

【问题讨论】:

    标签: python apache-spark pyspark spark-dataframe


    【解决方案1】:
    from pyspark.sql.functions import concat_ws, collect_list, concat, coalesce, lit
    
    #sample data
    df = sc.parallelize([
        [1, 'a', None, None],
        [1, None, 'b', None],
        [1, None, None, 'c'],
        [2, 'd', None, None],
        [2, None, 'e', None],
        [2, None, None, 'f']]).\
        toDF(('x', 'y', 'z', 'w'))
    df.show()
    
    result_df = df.groupby("x").\
                   agg(concat_ws(' ', collect_list(concat(*[coalesce(c, lit("")) for c in df.columns[1:]]))).
                       alias('result'))
    result_df.show()
    

    输出是:

    +---+------+
    |  x|result|
    +---+------+
    |  1| a b c|
    |  2| d e f|
    +---+------+
    

    示例输入:

    +---+----+----+----+
    |  x|   y|   z|   w|
    +---+----+----+----+
    |  1|   a|null|null|
    |  1|null|   b|null|
    |  1|null|null|   c|
    |  2|   d|null|null|
    |  2|null|   e|null|
    |  2|null|null|   f|
    +---+----+----+----+
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-11-23
      • 1970-01-01
      • 2016-09-12
      • 2019-03-07
      • 1970-01-01
      相关资源
      最近更新 更多