【问题标题】:Pyspark - Joins _ duplicate columnsPyspark - 加入_重复的列
【发布时间】:2020-09-04 13:53:11
【问题描述】:

我有 3 个数据框。 他们每个人都有如下所示的列:

我正在使用以下代码加入他们:

cond = [df1.col8_S1 == df2.col8_S1, df1.col8_S2 == df2.col8_S2]
df = df1.join(df2,cond,how ='inner').drop('df1.col8_S1','df1.col8_S2')
cond = [df.col8_S1 == df3.col8_S1, df.col8_S2 == df3.col8_S2]
df4 = df.join(df3,cond,how ='inner').drop('df3.col8_S1','df3.col8_S2')

我正在将数据帧写入 csv 文件;但是,由于它们从 col1 到 col7 具有相同的列,因此由于列重复,写入失败。如何在不指定名称的情况下删除重复的列。

【问题讨论】:

    标签: dataframe join pyspark duplicates


    【解决方案1】:

    只需使用列名进行连接,而不是显式使用相等操作。

    cond = ['col8_S1', 'col8_S2']
    df = df1.join(df2, cond, how ='inner')
    cond = ['col8_S1', 'col8_S2']
    df4 = df.join(df3, cond, how ='inner')
    

    【讨论】:

    • 它仍然没有消除重复的列
    猜你喜欢
    • 2022-01-06
    • 2021-11-04
    • 2019-02-09
    • 2020-03-27
    • 1970-01-01
    • 1970-01-01
    • 2022-01-26
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多