【发布时间】:2020-12-26 09:00:58
【问题描述】:
我正在尝试将四列(QBR、码、达阵和拦截)连接或组合成一列,并使用 pyspark 中的 f 等 sql 函数按球衣号码对它们进行分组。下面列出的是我尝试使用的编码、实际数据以及我预期的数据结果。
import pyspark.sql.functions as f
from pyspark.sql.functions import concat, lit, col
df = df.groupby('Jersey Number).withColumn("joined", f.concat(f.col('QBR'), f.lit(','), f.col('Yards'), f.lit(','), f.col('Touchdowns'), f.lit(','), f.col('Interceptions'))
Name Jersey Number QBR Yards Touchdowns Interceptions Fumbles
Kyler Murray 1 123.5 4120 40 6
Drew Brees 9 132.1 4500 52 12
Philip Rivers 17 120.4 3800 27 5
Andy Dalton 14 105.6 3650 22 7
Jersey Number Stats
1 123.5, 4120, 40, 6
9 132.1, 4500, 52, 12
14 105.6, 3650, 22, 7
17 120.4, 3800, 27, 5
【问题讨论】:
标签: python pyspark apache-spark-sql google-colaboratory