【问题标题】:How do I sum up values in a column into groups that match a given condition by date in pandas?如何将列中的值汇总为与熊猫中按日期匹配的给定条件的组?
【发布时间】:2021-06-17 12:20:03
【问题描述】:

我有一个这样的年龄组数据框

    Date            AgeGroup            Quantity
1   2020-12-08      18 - 29             1
2   2020-12-08      30 - 49             4
3   2020-12-08      50 - 54             0
4   2020-12-08      55 - 59             1
5   2020-12-08      60 - 64             1
6   2020-12-08      65 - 69             0
7   2020-12-08      70 - 74             3
8   2020-12-08      75 - 79             2
9   2020-12-08      80+                 1
....
10  2020-12-09      18 - 29             0
11  2020-12-09      30 - 49             2
12  2020-12-09      50 - 54             1
13  2020-12-09      55 - 59             2
14  2020-12-09      60 - 64             3
15  2020-12-09      65 - 69             0
16  2020-12-09      70 - 74             1
17  2020-12-09      75 - 79             1
18  2020-12-09      80+                 1

我想按三个更广泛的年龄组总结每个日期的数字,如下所示:

1   2020-12-08      18 - 59             6
2   2020-12-08      60+                 7
3   2020-12-08      75+                 3
...
4   2020-12-09      18 - 59             5
5   2020-12-09      60+                 5
6   2020-12-09      75+                 2

(第一个年龄组为 59 岁以下,第二个为 60 岁以上,第三个为 75 岁以上)。

我尝试了以下方法:

df1 = df.loc[(df['AgeGroup'] == '18 - 29') | (df['AgeGroup'] == '30 - 49') | (df['AgeGroup'] == '50 - 54') | (df['AgeGroup'] == '55 - 59') , 'Quantity'].sum()

但是,这缺少按日期细分,因为它只是为我提供了所有日期的这些年龄组的总和。

我也试过这个。

df.groupby(['Date', 'AgeGroup'])['Quantity'].sum()
Date        AgeGroup         
2020-12-08  18 - 29               1
            30 - 49               4
            50 - 54               0
            55 - 59               2
            60 - 64               1
            65 - 69               0
            70 - 74               3
            75 - 79               2

2020-12-09  18 - 29               0
            30 - 49               2
            50 - 54               1
            55 - 59               2
            60 - 64               3
            65 - 69               0
            70 - 74               1
            75 - 79               1

我仍然不知道如何在日期内组合这些年龄组。谢谢你的任何想法。

【问题讨论】:

    标签: python-3.x pandas dataframe sum pandas-groupby


    【解决方案1】:

    您可以通过Series.str.extract获取第一个数值,通过60比较并通过np.where设置为2组:

    m = df['AgeGroup'].str.extract('(\d+)', expand=False).astype(int) < 60
    df['AgeGroup'] = np.where(m, '18 - 59', '60+')
    
    df1 = df.groupby(['Date', 'AgeGroup'])['Quantity'].sum()
    print (df1)
    Date        AgeGroup
    2020-12-08  18 - 59     7
                60+         6
    2020-12-09  18 - 59     5
                60+         5
    Name: Quantity, dtype: int64
    

    【讨论】:

    • 谢谢。这对我来说不太管用,因为这些是 NHS 号码。有时,在我的数据集中,年龄组被定义为“30 至 39”或“60+”或“60+”。但我会查找 np.where..
    • @bluetail - 那么可以提取第一个数值吗?答案已编辑。
    • 最后通过将 groupbys 与您建议分成 3 组的条件相结合。
    猜你喜欢
    • 2015-03-29
    • 1970-01-01
    • 2021-05-19
    • 2021-01-06
    • 1970-01-01
    • 2023-01-03
    • 1970-01-01
    • 2018-09-07
    • 2021-02-09
    相关资源
    最近更新 更多