【问题标题】:Assign numeric values for multiple columns based on multiple conditions in pandas DataFrame根据 pandas DataFrame 中的多个条件为多列分配数值
【发布时间】:2022-03-17 09:31:43
【问题描述】:

我有一个包含几十列的 DataFrame。

Therapy area    Procedures1 Procedures2 Procedures3
Oncology        450         450         2345
Oncology        367         367         415
Oncology        152         152         4945
Oncology        876         876         345
Oncology        1098        1098        12
Oncology        1348        1348        234
Nononcology     225         225         345
Nononcology     300         300         44
Nononcology     267         267         45
Nononcology     90          90          4567

我想将所有Procedure 列中的数值更改为存储桶。

对于一列,它将类似于

def hello(x):
    if x['Therapy area'] == 'Oncology' and x['Procedures1'] < 200: return int(1)
    if x['Therapy area'] == 'Oncology' and x['Procedures1'] in range (200, 500): return 2
    if x['Therapy area'] == 'Oncology' and x['Procedures1'] in range (500, 1000): return 3
    if x['Therapy area'] == 'Oncology' and x['Procedures1'] > 1000: return 4
    if x['Therapy area'] != 'Oncology' and x['Procedures1'] < 200: return 11
    if x['Therapy area'] != 'Oncology' and x['Procedures1'] in range (200, 500): return 22
    if x['Therapy area'] != 'Oncology' and x['Procedures1'] in range (500, 1000): return 33
    if x['Therapy area'] != 'Oncology' and x['Procedures1'] > 1000: return 44  
test['Procedures1'] = test.apply(hello, axis=1)

对具有不同列名(不是Procedures1、Procedures2、'Procedures3` 等)的几十个列应用此方法的最有效方法是什么?

将 cut 与特定的垃圾箱一起使用时,我收到错误:

ValueError: bins 必须单调增加。

既然我可以有不同的值,我该如何用逻辑运算而不是 bin 来解决这个问题?

此外,值可能会因“治疗领域”列而异,例如非肿瘤学为 11、22、33、44,肿瘤学为 1、2、3、4。

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    你可以applypd.cut到相关栏目:

    cols = ['Procedures1', 'Procedures2']
    df[cols] = df[cols].apply(lambda col: pd.cut(col, [0,200,500,1000, col.max()], labels=[1,2,3,4]))
    

    输出:

      Therapy_area Procedures1 Procedures2
    0     Oncology           2           2
    1     Oncology           2           2
    2     Oncology           1           1
    3     Oncology           3           3
    4     Oncology           4           4
    5     Oncology           4           4
    6  Nononcology           2           2
    7  Nononcology           2           2
    8  Nononcology           2           2
    9  Nononcology           1           1
    

    你也可以使用np.select:

    def encoding(col, labels):
        return np.select([col<200, col.between(200,500), col.between(500,1000), col>1000], labels, 0)
    
    onc_labels = [1,2,3,4]
    nonc_labels = [11,22,33,44]
    msk = df['Therapy_area'] == 'Oncology'
    
    df[cols] = pd.concat((df.loc[msk, cols].apply(encoding, args=(onc_labels,)), df.loc[msk, cols].apply(encoding, args=(nonc_labels,)))).reset_index(drop=True)
    

    输出:

      Therapy_area  Procedures1  Procedures2  Procedures3
    0     Oncology            2            2            4
    1     Oncology            2            2            2
    2     Oncology            1            1            4
    3     Oncology            3            3            2
    4     Oncology            4            4            1
    5     Oncology            4            4            2
    6  Nononcology           22           22           44
    7  Nononcology           22           22           22
    8  Nononcology           11           11           44
    9  Nononcology           33           33           22
    

    【讨论】:

      猜你喜欢
      • 2018-04-30
      • 2018-04-28
      • 2021-04-30
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-03-19
      • 1970-01-01
      相关资源
      最近更新 更多