【发布时间】:2020-09-04 13:53:11
【问题描述】:
我有 3 个数据框。 他们每个人都有如下所示的列:
我正在使用以下代码加入他们:
cond = [df1.col8_S1 == df2.col8_S1, df1.col8_S2 == df2.col8_S2]
df = df1.join(df2,cond,how ='inner').drop('df1.col8_S1','df1.col8_S2')
cond = [df.col8_S1 == df3.col8_S1, df.col8_S2 == df3.col8_S2]
df4 = df.join(df3,cond,how ='inner').drop('df3.col8_S1','df3.col8_S2')
我正在将数据帧写入 csv 文件;但是,由于它们从 col1 到 col7 具有相同的列,因此由于列重复,写入失败。如何在不指定名称的情况下删除重复的列。
【问题讨论】:
标签: dataframe join pyspark duplicates