【问题标题】:Pandas merging dataframes熊猫合并数据框
【发布时间】:2017-06-29 11:59:15
【问题描述】:

我有几个要合并的数据框,但问题是它们没有相同的列,我只想合并特定的行。我将展示一个示例,这样会更容易:

MAIN_DF 我希望所有的都合并到它:

key    A    B    C
0001   1    0    0
0002   1    1    1
0003   0    0    1

DF_1:

key    A    B    C   D
0001   1    0    0   1
0003   0    0    1   0
0004   1    1    1   1

DF_2:

key    C    D    E   F
0004   1    1    0   1
0005   0    0    1   0
0006   1    1    1   1

所以我想将它全部合并到 MAIN_DF,所以 MAIN_DF 将是:

key    A    B    C    D    E   F
0001   1    0    0    1    0   0
0002   1    1    1    0    0   0
0003   0    0    1    0    0   0
0004   0    0    0    1    0   1
0005   0    0    0    0    1   0
0006   0    0    0    1    1   1

查看列是否已更新并添加了新行。

是否可以在不使用长而慢的循环和 if 语句的情况下使用 pandas 来做到这一点?

谢谢

【问题讨论】:

  • 输出中带有0004 的行是否正确?
  • 您的左下 3 个角单元格应从上到下读取 [[1,1,1], [0,0,0], [0,0,1]]。

标签: python pandas dataframe


【解决方案1】:

我觉得你需要DataFrame.combine_first:

MAIN_DF = MAIN_DF.set_index('key')
DF_1 = DF_1.set_index('key')
DF_2 = DF_2.set_index('key')

df = MAIN_DF.combine_first(DF_1).combine_first(DF_2).fillna(0).astype(int).reset_index()
print (df)
    key  A  B  C  D  E  F
0  0001  1  0  0  1  0  0
1  0002  1  1  1  0  0  0
2  0003  0  0  1  0  0  0
3  0004  1  1  1  1  0  1
4  0005  0  0  0  0  1  0
5  0006  0  0  1  1  1  1

【讨论】:

    【解决方案2】:

    这是使用groupby 的方法。

    import pandas as pd 
    import numpy as np
    
    df1 = pd.DataFrame([[1, 0, 0],
                        [1, 1, 1],
                        [0, 0, 1]],    columns=['a', 'b', 'c'],      index=[1, 2, 3])
    df2 = pd.DataFrame([[1, 0, 0, 1],
                        [0, 0, 1, 0],
                        [1, 1, 1, 1]], columns=['a', 'b', 'c', 'd'], index=[1, 3, 4])
    df3 = pd.DataFrame([[1, 1, 0, 1],
                        [0, 0, 1, 0],
                        [1, 1, 1, 1]], columns=['c', 'd', 'e', 'f'], index=[4, 5, 6])
    
    # combine the first and second df
    df4 = pd.concat([df1, df2])
    grouped = df4.groupby(level=0)
    df5 = grouped.first()
    
    # combine (first and second combined), with the third
    df6 = pd.concat([df5, df3])
    grouped = df6.groupby(level=0)
    df7 = grouped.first()
    
    # fill na values with 0
    df7.fillna('0', inplace=True)
    
    print(df)
    
        a   b   c   d   e   f
    1   1   0   0   1   0   0
    2   1   1   1   0   0   0
    3   0   0   1   0   0   0
    4   1   1   1   1   0   1
    5   0   0   0   0   1   0
    6   0   0   1   1   1   1
    

    【讨论】:

      【解决方案3】:

      您可以使用 concat 水平连接任意数量的数据帧:

      import pandas as pd
      df = pd.concat([df1,df2], axis=1, verify_integrity=True)
      

      “verify_integrity”参数检查重复项。

      点击这里了解更多关于merge, join and concatenate

      【讨论】:

      • 是的,但请注意我不想重复行
      猜你喜欢
      • 2013-09-26
      • 2018-02-02
      • 2017-06-11
      • 2016-01-01
      • 2016-10-31
      • 2014-07-02
      • 1970-01-01
      相关资源
      最近更新 更多