【问题标题】:DataFrame GroupBy multi level selectionDataFrame GroupBy 多级选择
【发布时间】:2019-02-12 01:10:21
【问题描述】:

我正在尝试使用 pandas 来解决我用纯 python 完成的问题,但不知道 DataFrame groupby 的最佳实践。

我想为每个邮政编码选择处方最多的药物(占该邮政编码中所有药物的百分比)。 如果两种药物的处方数量相同,我想选择“按字母顺序排列第一个”的药物:

import pandas as pd

drugs_prescriptions = pd.DataFrame({'PostCode': ['P1', 'P1', 'P1', 'P2', 'P2', 'P3'],
                    'Drug': ['D1', 'D2', 'D1', 'D2', 'D1', 'D2'],
                    'Quantity': [3, 6, 5, 7, 7, 8]})

    Drug    PostCode    Quantity
# 0 D1      P1          3
# 1 D2      P1          6
# 2 D1      P1          5
# 3 D2      P2          7
# 4 D1      P2          7
# 5 D2      P3          8

#This should be the RESULT:
# postCode, drug with highest quantity, percentage of all drugs per post code
# (post code P2 has two drugs with the same quantity, alphabetically first one is selected
# [('P1', 'D1', 0.57),
# ('P2', 'D1', 0.50),
# ('P3', 'D2', 1)]

我已按邮政编码、药物进行分组,但在选择行时遇到问题(应用 lambda)。

durg_qualtity_per_post_code = drugs_prescriptions.groupby(['PostCode', 'Drug']).agg('sum')

每个邮政编码销售的所有药物,我打算将这个与应用或转换一个以前的数据集一起使用:

all_by_post_code = drugs_prescriptions.groupby(['PostCode'])['Quantity'].sum()

我不确定如何选择每个邮政编码的药物最大数量行,如果两种药物的数量相同,则应选择具有第一个字母顺序的药物(邮政编码 P2 为 D1)。

我想做这样的事情:

durg_qualtity_per_post_code [durg_qualtity_per_post_code .apply(lambda x: int(x['Quantity']) == max_items_by_post_code[x['post_code']], axis=1, reduce=True)]

更新:

# sort by PostCode, Drug
df = drugs_prescriptions.groupby(['PostCode', 'Drug']).agg('sum')
df = df.groupby(['PostCode']).apply(lambda x: x.sort_values(['Quantity', 'Drug'], ascending=[False, True]))

# select first value by PostCode
# reset index in order to have drug in the output as well
df.reset_index(level=[1], inplace=True)
df = df.groupby(['PostCode']).first()

# calculate percentage of total by PostCode
allQuantities = drugs_prescriptions.groupby(['PostCode']).agg('sum')
df['Quantity'] = df.apply(lambda row: row['Quantity']/allQuantities.loc[row.name], axis=1)

【问题讨论】:

    标签: python python-3.x pandas


    【解决方案1】:

    这是一种可能的解决方案,但感觉很尴尬且不符合 Python 风格。但它有效,cmets 在代码中。

    # setting string to integer
    df.Quantity = df.Quantity.astype('int')
    
    # create a mulitiindex
    df.set_index(['PostCode', 'Drug'], inplace=True)
    
    # use transform to divide the sum of the 'Drug' level by the 'PostCode' level
    df = df.groupby(level=[0,1]).transform('sum') / df.groupby(level=0).transform('sum')
    
    # move 'Drug' out of the multi index to allow for sorting
    df.reset_index(level=[1], inplace=True)
    
    # Sort the 'Quantity' descending order, and the 'Drug' in ascending order,
    # then we can select the first 'PostCode' for our result
    df.sort_values(['Quantity','Drug'], ascending=[False, True], inplace=True)
    
    df.groupby('PostCode').first()
    
               Drug Quantity
    PostCode        
    P1          D1  0.571429
    P2          D1  0.500000
    P3          D2  1.000000
    

    【讨论】:

    • 谢谢,有很多资料要研究。真的很感激。
    • @user007 这是一个很好的问题,我自己也学到了一些东西。谢谢。
    猜你喜欢
    • 2016-07-16
    • 2013-02-19
    • 2019-05-06
    • 2021-11-03
    • 2021-06-18
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多