【发布时间】:2021-03-21 02:13:09
【问题描述】:
我有产品、品牌和百分比列。我想计算与当前行具有不同品牌的行以及与当前行具有相同品牌的行的百分比列的总和。如何在 PySpark 中或使用 spark.sql 来完成?
样本数据:
df = pd.DataFrame({'a': ['a1','a2','a3','a4','a5','a6'],
'brand':['b1','b2','b1', 'b3', 'b2','b1'],
'pct': [40, 30, 10, 8,7,5]})
df = spark.createDataFrame(df)
我在寻找什么:
product brand pct pct_same_brand pct_different_brand
a1 b1 40 null null
a2 b2 30 null 40
a3 b1 10 40 30
a4 b3 8 null 80
a5 b2 7 30 58
a6 b1 5 50 45
这是我尝试过的:
df.createOrReplaceTempView('tmp')
spark.sql("""
select *, sum(pct * (select case when n1.brand = n2.brand then 1 else 0 end
from tmp n1)) over(order by pct desc rows between UNBOUNDED PRECEDING and 1
preceding)
from tmp n2
""").show()
【问题讨论】:
标签: apache-spark pyspark hive apache-spark-sql