【问题标题】:Python - replace values in dataframe based on another dataframe matchPython - 根据另一个数据框匹配替换数据框中的值
【发布时间】:2019-03-23 04:16:16
【问题描述】:

假设我有 2 个名为 A 和 B 的 Python 数据框2,如下所示。 如何根据 B 中的列 ID 和月份的匹配替换数据框 A 中的列值? 有什么想法吗?

谢谢

数据框 A:

ID  Month   City    Brand   Value
1   1   London  Unilever    100
1   2   London  Unilever    120
1   3   London  Unilever    150
1   4   London  Unilever    140
2   1   NY  JP Morgan   90
2   2   NY  JP Morgan   105
2   3   NY  JP Morgan   100
2   4   NY  JP Morgan   140
3   1   Paris   Loreal  60
3   2   Paris   Loreal  75
3   3   Paris   Loreal  65
3   4   Paris   Loreal  80
4   1   Tokyo   Sony    100
4   2   Tokyo   Sony    90
4   3   Tokyo   Sony    85
4   4   Tokyo   Sony    80

数据框 B:

ID  Month   Value
2   1   100
3   3   80

【问题讨论】:

    标签: python pandas numpy replace


    【解决方案1】:

    合并它们,然后删除未使用的字段:

    C = pd.merge(A[['ID', 'Month', 'City', 'Brand']],B, on=['ID', 'Month'])
    C = C[['ID', 'Month', 'City', 'Brand', 'Value']]
    

    这应该可以工作

    【讨论】:

    • 你测试了吗?
    • 嗨,这很酷,但它只将数据帧与过滤合并。我的问题更多是关于在两列中都存在匹配时替换 Value 列数据的行。转换为 SQL 时,相当于在 2 列上连接更新语句。有什么猜测吗?干杯
    • @jezrael,是的,但我看到了问题。
    【解决方案2】:

    将merge 与左连接一起使用,并将缺失值替换为fillna 的原始值:

    df = df1.merge(df2, on=['ID', 'Month'], how='left', suffixes=('_',''))
    df['Value'] = df['Value'].fillna(df['Value_']).astype(int)
    df = df.drop('Value_', axis=1)
    print (df)
        ID  Month    City      Brand  Value
    0    1      1  London   Unilever    100
    1    1      2  London   Unilever    120
    2    1      3  London   Unilever    150
    3    1      4  London   Unilever    140
    4    2      1      NY  JP Morgan    100
    5    2      2      NY  JP Morgan    105
    6    2      3      NY  JP Morgan    100
    7    2      4      NY  JP Morgan    140
    8    3      1   Paris     Loreal     60
    9    3      2   Paris     Loreal     75
    10   3      3   Paris     Loreal     80
    11   3      4   Paris     Loreal     80
    12   4      1   Tokyo       Sony    100
    13   4      2   Tokyo       Sony     90
    14   4      3   Tokyo       Sony     85
    15   4      4   Tokyo       Sony     80
    

    【讨论】:

    • 谢谢。我收到一个错误,可能与“Value_”部分有关。键错误:“值_”。关于原因的任何想法?谢谢
    • print (df1.columns) 是什么? print (df2.columns) 因为Value 是样本数据的最后一列,Value_ 也是。
    • 是的,我知道 print 语句的目的,只是想知道为什么我无法得到与您相同的结果?
    • 所以因为在样本数据中是名为Value的列,所以keyerror意味着在实际数据中没有列Value。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-05-06
    • 2020-11-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多