【问题标题】:How can I add a new column to pandas dataframe based on groupby function output?如何根据 groupby 函数输出向 pandas 数据框添加新列?
【发布时间】:2017-06-29 18:40:20
【问题描述】:

我有一个包含 500,000 行的 dataframe1。我想通过在包含配置的 dataframe2 中查找型号来填充配置列。

数据框1:

 Model                 Date     Status   Configuration
 A4                    10/2014  Inop      
 A4                    11/2014  Op              
 A4                    11/2014  Op                                     
 G5                    10/2014  Inop                                   
 G5                    11/2014  Inop                                   
 G5                    11/2014  Op                                     
 G8                    10/2014  Op                                     
 G8                    11/2014  Op                                     
 G8                    11/2014  Op                                     
 G8                    10/2014  Inop                                   
 Z2                    11/2014  Op                                     
 Z2                    11/2014  Op                                     

数据框2:

 Model              Configuration  
 A4                 ICS   
 G5                 PCS  
 G8                 ICS    
 Z2                 1/2 ICS   

我目前正在运行的代码:

for Model, group in dataframe1.groupby('Model'):
    #gets configuration from dataframe2 
    config = get_configuration(Model)
    #attempt to assign configuration to all columns with that model number in dataframe1
    dataframe1['Config'] = con

此代码返回:

此代码按模型对 dataframe1 进行分组并成功获取每个组的配置,但我无法将该配置应用于 dataframe1 中的新行以获得以下结果:

 Model                 Date     Status   Configuration
 A4                    10/2014  Inop     ICS   
 A4                    11/2014  Op       ICS     
 A4                    11/2014  Op       ICS     
 G5                    10/2014  Inop     PCS   
 G5                    11/2014  Inop     PCS  
 G5                    11/2014  Op       PCS
 G8                    10/2014  Op       ICS 
 G8                    11/2014  Op       ICS      
 G8                    11/2014  Op       ICS      
 G8                    10/2014  Inop     ICS     
 Z2                    11/2014  Op       1/2 ICS 
 Z2                    11/2014  Op       1/2 ICS

【问题讨论】:

标签: python pandas dataframe jupyter


【解决方案1】:

使用map

Dataframe1['Config'] = Dataframe1['Model'].map(Dataframe2.set_index('Model').Config)
Dataframe1

   Model     Date Status   Config
0     A4  10/2014   Inop      ICS
1     A4  11/2014     Op      ICS
2     A4  11/2014     Op      ICS
3     G5  10/2014   Inop  Non ICS
4     G5  11/2014   Inop  Non ICS
5     G5  11/2014     Op  Non ICS
6     G8  10/2014     Op      ICS
7     G8  11/2014     Op      ICS
8     G8  11/2014     Op      ICS
9     G8  10/2014   Inop      ICS
10    Z2  11/2014     Op  1/2 ICS
11    Z2  11/2014     Op  1/2 ICS

【讨论】:

    【解决方案2】:

    试试pd.merge

    Dataframe1.merge(Dataframe2,left_on='Model',right_on='Model',how='left')         
    

    【讨论】:

    • 这也是一个很好的解决方案 :-)... 如果列名相同,则不需要 right_on 或 left_on。你可以使用on
    • @piRSquared 效率,你的更好~
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-06-20
    • 2021-03-02
    • 2020-02-24
    相关资源
    最近更新 更多