【问题标题】:GroupBy items and count item every hour in Pandas在 Pandas 中每小时按项目分组并计数项目
【发布时间】:2018-11-27 14:13:48
【问题描述】:

我有以下数据集,以“日期时间”对象为索引

index                                Item

2016-10-30 09:58:11                 Bread
2016-10-30 10:05:34          Scandinavian
2016-10-30 10:05:34          Scandinavian
2016-10-30 10:07:57         Hot chocolate
2016-10-30 10:07:57                   Jam
2016-10-30 10:07:57               Cookies
2016-10-30 10:19:12                Pastry
2016-10-30 10:19:12                Coffee
2016-10-30 10:19:12                   Tea
2016-10-30 10:20:51                Pastry
2016-10-30 10:20:51                 Bread
2016-10-30 10:21:59                 Bread
2016-10-30 10:21:59                Muffin

作为 Pandas 的新手,我对如何按数据框进行分组有点迷茫。我需要两件事 1) 每小时的项目计数,例如每小时“面包”的总计数

类似下面的东西

index           item          count

 2016-10-30 09:00:00   Bread   3
 2016-10-30 10:00:00  Coffee  10
 2016-10-30 11:00:00   Toast   1

然后是一天 24 小时内的项目总数

index          item  count

 2016-10-30    Bread  13
 2016-10-30   Coffee  1200
 2016-10-30    Toast  19

可能是两个独立的操作?

【问题讨论】:

    标签: python pandas pandas-groupby


    【解决方案1】:

    获取DatetimeIndex.floor并通过GroupBy.size聚合:

    print (type(df))
    <class 'pandas.core.frame.DataFrame'>
    
    dates = df.rename_axis('Dates').index.floor('H')
    df1 = df.groupby([dates,'Item']).size().reset_index(name='count')
    print (df1)
                    Dates           Item  count
    0 2016-10-30 09:00:00          Bread      1
    1 2016-10-30 10:00:00          Bread      2
    2 2016-10-30 10:00:00         Coffee      1
    3 2016-10-30 10:00:00        Cookies      1
    4 2016-10-30 10:00:00  Hot chocolate      1
    5 2016-10-30 10:00:00            Jam      1
    6 2016-10-30 10:00:00         Muffin      1
    7 2016-10-30 10:00:00         Pastry      2
    8 2016-10-30 10:00:00   Scandinavian      2
    9 2016-10-30 10:00:00            Tea      1
    

    dates = df.rename_axis('Dates').index.floor('24H')
    df2 = df.groupby([dates,'Item']).size().reset_index(name='count')
    print (df2)
           Dates           Item  count
    0 2016-10-30          Bread      3
    1 2016-10-30         Coffee      1
    2 2016-10-30        Cookies      1
    3 2016-10-30  Hot chocolate      1
    4 2016-10-30            Jam      1
    5 2016-10-30         Muffin      1
    6 2016-10-30         Pastry      2
    7 2016-10-30   Scandinavian      2
    8 2016-10-30            Tea      1
    

    如果Series:

    print (type(s))
    <class 'pandas.core.series.Series'>
    
    dates = s.rename_axis('Dates').index.floor('24H')
    df2 = s.groupby([dates,s]).size().reset_index(name='count')
    

    【讨论】:

    • 编辑问题,注意索引是日期时间对象
    • @Wajih - 是的,我的解决方案仅适用于 DatetimeIndex
    • 正在尝试。给我几分钟
    • 超级!感谢您的解决方案。效果很棒!
    猜你喜欢
    • 1970-01-01
    • 2011-08-07
    • 1970-01-01
    • 1970-01-01
    • 2017-10-18
    • 1970-01-01
    • 2020-11-16
    • 1970-01-01
    • 2021-10-13
    相关资源
    最近更新 更多