【问题标题】:Combining 4 Columns and Grouping By One Column using Pyspark SQL Functions使用 Pyspark SQL 函数组合 4 列并按一列分组
【发布时间】:2020-12-26 09:00:58
【问题描述】:

我正在尝试将四列(QBR、码、达阵和拦截)连接或组合成一列,并使用 pyspark 中的 f 等 sql 函数按球衣号码对它们进行分组。下面列出的是我尝试使用的编码、实际数据以及我预期的数据结果。

import pyspark.sql.functions as f
from pyspark.sql.functions import concat, lit, col
df = df.groupby('Jersey Number).withColumn("joined", f.concat(f.col('QBR'), f.lit(','), f.col('Yards'), f.lit(','), f.col('Touchdowns'), f.lit(','), f.col('Interceptions'))
Name           Jersey Number      QBR        Yards    Touchdowns     Interceptions Fumbles
Kyler Murray       1              123.5      4120      40             6
Drew Brees         9              132.1      4500      52             12
Philip Rivers      17             120.4      3800      27             5
Andy Dalton        14             105.6      3650      22             7



Jersey Number   Stats       
    1           123.5, 4120, 40, 6
    9           132.1, 4500, 52, 12
    14          105.6, 3650, 22, 7
    17          120.4, 3800, 27, 5

【问题讨论】:

    标签: python pyspark apache-spark-sql google-colaboratory


    【解决方案1】:

    尝试使用 concat_ws, flatten, collect_list(array(cols)) 函数。

    Example:

    df.show()
    #+-------------+-----+-----+-----+-------------+
    #|Jersey number|  QBR|yards|touch|intercepyions|
    #+-------------+-----+-----+-----+-------------+
    #|            1|123.5| 4120|   40|            6|
    #+-------------+-----+-----+-----+-------------+
    
    from pyspark.sql.functions import *
    
    df.groupBy("Jersey number").\
    agg(concat_ws(",",flatten(collect_list(array(*cols)))).alias("Stats")).\
    show(10,False)
    #+-------------+---------------------+
    #|Jersey number|Stats                |
    #+-------------+---------------------+
    #|1            |123.5,4120.0,40.0,6.0|
    #+-------------+---------------------+
    
    df.groupBy("Jersey number").agg(array_join(flatten(collect_list(array(*cols))),',').alias("stats")).show(10,False)
    #+-------------+---------------------+
    #|Jersey number|stats                |
    #+-------------+---------------------+
    #|1            |123.5,4120.0,40.0,6.0|
    #+-------------+---------------------+
    

    import as f:

    from pyspark.sql import functions as f
    
    cols = ['QBR', 'yards', 'touch', 'intercepyions']
    
    df.groupBy("Jersey number").agg(f.concat_ws(",",f.flatten(f.collect_list(f.array(*cols)))).alias("Stats")).show(10,False)
    
    #or using array_join
    df.groupBy("Jersey number").agg(f.array_join(f.flatten(f.collect_list(f.array(*cols))),',').alias("stats")).show(10,False)
    

    【讨论】:

    • 你知道f函数是怎么做的吗?我需要学习如何使用 f 函数。
    • 对不起最后一个问题。如果数据中有其他列,但您只想使用为此函数列出的五列怎么办?
    • 定义一个列表,其中所需的列为cols=['col1','col2'...etc],然后在您的聚合中使用cols。
    • 我曾尝试在数据中这样做,但它似乎不起作用。你能给我举个例子吗?
    • 检查我的更新答案,包括定义列表cols = ['QBR', 'yards', 'touch', 'intercepyions']
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-06-17
    • 1970-01-01
    • 2018-03-02
    • 2018-11-08
    • 1970-01-01
    相关资源
    最近更新 更多