【问题标题】:Duplicate columns when pandas dataframes are merged合并熊猫数据框时重复列
【发布时间】:2020-09-10 11:18:07
【问题描述】:

我想合并 df1 和 df2。我在合并 df1 和 df2 时遇到的当前问题是它会产生重复的“Fluc”列。数据框必须合并on='Horse'。

数据框代码:

cols1 = ['Race', 'Horse', 'Fluc 1', 'Fluc 2','Bookmaker', 'Odds']
df1 = pd.DataFrame(data=data, columns=cols1)
cols2 = ['Race', 'Horse', 'Fluc 1', 'Fluc 2', 'Bookmaker', 'AvgOdds']
df2 = pd.DataFrame(data=data, columns=cols2)
df3 = df2.groupby(by='Horse', sort=False).mean()
df3 = df3.reset_index()
df4 = round(df3,2)
dfmerge = pd.merge(df1,df4,on='Horse',how='inner')

df1 的输出:

              Race           Horse  Fluc 1  Fluc 2      Bookmaker   Odds
0       Ipswich R1  Battle Through     4.2    4.22        BetEasy   4.20
1       Ipswich R1  Battle Through     4.2    4.22           Neds   4.20
2       Ipswich R1  Battle Through     4.2    4.22      Sportsbet   4.20
3       Ipswich R1  Battle Through     4.2    4.22  SportsBetting   4.45
4       Ipswich R1  Battle Through     4.2    4.22         Bet365   4.20

df4 的输出:

              Race           Horse  Fluc 1  Fluc 2      Bookmaker  AvgOdds
0       Ipswich R1  Battle Through     4.2    4.22        BetEasy     4.20
1       Ipswich R1  Battle Through     4.2    4.22           Neds     4.20
2       Ipswich R1  Battle Through     4.2    4.22      Sportsbet     4.20
3       Ipswich R1  Battle Through     4.2    4.22  SportsBetting     4.45
4       Ipswich R1  Battle Through     4.2    4.22         Bet365     4.20

dfmerge 的输出:

              Race           Horse  Fluc 1_x  Fluc 2_x      Bookmaker  Odds  Fluc 1_y  Fluc 2_y  AvgOdds
0       Ipswich R1  Battle Through      8.34      8.38           Neds   8.5      8.34      8.38     8.65
1       Ipswich R1  Battle Through      8.34      8.38      Sportsbet   8.0      8.34      8.38     8.65
2       Ipswich R1  Battle Through      8.34      8.38  SportsBetting   9.1      8.34      8.38     8.65
3       Ipswich R1  Battle Through      8.34      8.38         Bet365   9.0      8.34      8.38     8.65
4       Ipswich R1      Simply Fly      1.89      1.87           Neds   1.8      1.89      1.87     1.84

dfmerge 的期望输出:

              Race           Horse  Fluc 1  Fluc 2      Bookmaker   Odds    AvgOdds
0       Ipswich R1  Battle Through     4.2    4.22        BetEasy   4.20    4.2
1       Ipswich R1  Battle Through     4.2    4.22           Neds   4.20    4.2
2       Ipswich R1  Battle Through     4.2    4.22      Sportsbet   4.20    4.2
3       Ipswich R1  Battle Through     4.2    4.22  SportsBetting   4.45    4.2
4       Ipswich R1  Battle Through     4.2    4.22         Bet365   4.20    4.2

【问题讨论】:

  • 使用suffixes参数为重复的列添加后缀并根据后缀删除列
  • 您好,当您合并 df1 和 df4 时,您应该向我们展示 df4 而不是 df2 的输出。将后缀(_x 和 _y)添加到两个数据帧中的列是合并函数的常见行为。
  • 那么为什么不为'Bookmaker'、'Horse'等提供(_x和_y)列
  • 您是否只是想将 AvgOdds 列从 df2 引入 df1?如果是这种情况,您是否尝试过:how = 'left'?
  • 我需要合并 Fluc 列,以免重复。这是主要问题

标签: python pandas dataframe


【解决方案1】:

试试这个

dfmerge = pd.merge(df1, df4, on=['Race', 'Horse', 'Fluc 1', 'Fluc 2', 'Bookmaker'], how='inner')
print(dfmerge)

输出:

         Race           Horse  Fluc 1  Fluc 2      Bookmaker  Odds  AvgOdds
0  Ipswich R1  Battle Through     4.2    4.22        BetEasy  4.20     4.20
1  Ipswich R1  Battle Through     4.2    4.22           Neds  4.20     4.20
2  Ipswich R1  Battle Through     4.2    4.22      Sportsbet  4.20     4.20
3  Ipswich R1  Battle Through     4.2    4.22  SportsBetting  4.45     4.45
4  Ipswich R1  Battle Through     4.2    4.22         Bet365  4.20     4.20

【讨论】:

    猜你喜欢
    • 2018-11-04
    • 1970-01-01
    • 2017-11-26
    • 1970-01-01
    • 2016-10-31
    • 1970-01-01
    • 2015-02-03
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多