【问题标题】:Matching the column names of two pandas data-frames in python匹配python中两个熊猫数据框的列名
【发布时间】:2018-01-23 14:46:56
【问题描述】:

我有两个名为 df1 和 df2 的 pandas 数据框,这样 `

df1: a      b    c    d
     1      2    3    4
     5      6    7    8

和

df2:      b    c   
          12   13   

我希望结果像

result:    b    c    
           2    3    
           6    7    

这里要注意a b c d是pandas dataframe中的列名。两个熊猫数据框的形状和值都不同。我想将df2的列名与df1的列名匹配,并选择df1的所有行,其标题与df2的列名匹配。df2仅用于选择维护所有行的df1 的特定列。我尝试了下面给出的一些代码,但这给了我一个空索引。

df1.columns.intersection(df2.columns)

上面的代码没有给出我的结果,因为它给出了没有值的索引标题。我想编写一个代码,我可以在其中将我的两个数据框作为输入,并比较列标题以供选择。我不必对列名进行硬编码。

【问题讨论】:

  • 看来您只需要df1[df1.columns.intersection(df2.columns)]。
  • 你有列名,你只需要用它来索引df1。
  • 或者,df1[df1.columns & df2.columns]

标签: python pandas pandas-groupby


【解决方案1】:

我相信你需要:

df = df1[df1.columns.intersection(df2.columns)]

或者像 @Zero 在 cmets 中指出的那样:

df = df1[df1.columns & df2.columns]

【讨论】:

  • 不需要concat,看看预期的结果;-)
【解决方案2】:

或者,使用reindex

In [594]: df1.reindex(columns=df2.columns)
Out[594]:
   b  c
0  2  3
1  6  7

也就是

In [595]: df1.reindex(df2.columns, axis=1)
Out[595]:
   b  c
0  2  3
1  6  7

【讨论】:

    【解决方案3】:

    或者交叉路口:

    df = df1[df1.columns.isin(df2.columns)]
    

    【讨论】:

      猜你喜欢
      • 2019-06-21
      • 1970-01-01
      • 1970-01-01
      • 2021-08-08
      • 2021-05-15
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多