【问题标题】:how to match two columns using one of the column as reference?如何使用其中一列作为参考来匹配两列?
【发布时间】:2016-12-23 21:37:55
【问题描述】:

我做了一些分析并找到了一个特定的模式,现在我正在尝试做一些预测。 我有一个数据集,可以预测童年时期发生给定事故次数的学生的评分。 我的预测矩阵看起来像这样:

   A
   injuries      ratings  
         0            5
         1            4.89
         2            4.34
         3            3.99 
         4            3.89
         5            3.77 

我的数据集如下所示:

B

siblings income injuries total_scoldings_from father
     3       12000   4             09
     4       34000   5             22
     1       23400   3             12
     3       24330   1              1
     0       12000   1             12 

现在我想创建一个列名predictions,它基本上匹配从A 到B 的条目并返回

siblings income injuries total_scoldings_from_father predictions
     3       12000   4             09                    3.89
     4       34000   5             22                    3.77
     1       23400   3             12                    3.99
     3       24330   1             1                     4.89
     0       12000   1             12                    4.89

请帮忙

还建议一个标题,因为我的标题缺少对将来参考的所有重要内容

【问题讨论】:

    标签: python pandas dataframe mapping multiple-columns


    【解决方案1】:

    如果映射的所有值都在 DataFrame A 中,您可以使用 map:

    B['predictions'] = B['injuries'].map(A.set_index('injuries')['ratings'])
    print (B)
       siblings  income  injuries  total_scoldings_from_father  predictions
    0         3   12000         4                            9         3.89
    1         4   34000         5                           22         3.77
    2         1   23400         3                           12         3.99
    3         3   24330         1                            1         4.89
    4         0   12000         1                           12         4.89
    

    merge 的另一个解决方案:

    C = pd.merge(B,A)
    print (C)
       siblings  income  injuries  total_scoldings_from_father  ratings
    0         3   12000         4                            9     3.89
    1         4   34000         5                           22     3.77
    2         1   23400         3                           12     3.99
    3         3   24330         1                            1     4.89
    4         0   12000         1                           12     4.89
    

    【讨论】:

    • 如果 B 的行数比 A 多得多,将合并工作,因为 A 只是一个参考矩阵?我的头衔也合适吗?
    • 是的,DataFrame A 用作参考,merge 解决方案有效,如果 A 只有两列。如果有多个列,使用子集C = pd.merge(B,A[['injuries','ratings']])
    猜你喜欢
    • 1970-01-01
    • 2022-10-01
    • 2015-02-11
    • 1970-01-01
    • 1970-01-01
    • 2016-04-11
    • 2019-04-02
    • 2021-11-04
    • 2023-04-08
    相关资源
    最近更新 更多