【问题标题】:How to group data on Pandas with multiple conditions?如何在多个条件下对 Pandas 上的数据进行分组?
【发布时间】:2018-01-25 21:27:37
【问题描述】:

这是我的桌子

timestamp        date month  day   hour   price
0  2017-01-01 00:00  01/01/2017   Jan  Sun  00:00  60.23
1  2017-01-01 01:00  01/01/2017   Jan  Sun  01:00  60.73
2  2017-01-01 02:00  01/01/2017   Jan  Sun  02:00  75.99
3  2017-01-01 03:00  01/01/2017   Jan  Sun  03:00  60.76
4  2017-01-01 04:00  01/01/2017   Jan  Sun  04:00  49.01

我每天都有一天 24 小时的数据,以及一整年的每个月的数据。

例如,我想对工作日和周末的每个季节的数据进行分组 Weekend_Winter = 11 月、12 月、1 月、2 月的所有周六和周日数据

这方面的新手,所以任何帮助都会很有用

【问题讨论】:

  • @user3256363 组是什么意思?将它们拆分为不同的 DataFrame?做一个groupby?用所述组的名称分配一个新列?

标签: python pandas


【解决方案1】:

以下解决方案与@jezrael 略有不同,因为明确定义了季节和工作日。

import pandas as pd

df = pd.DataFrame([['2017-01-01 00:00', '01/01/2017', 'Jan', 'Mon', '00:00', 60.23],
                   ['2017-01-01 01:00', '01/01/2017', 'Jan', 'Sat', '01:00', 60.73],
                   ['2017-01-01 02:00', '01/01/2017', 'May', 'Tue', '02:00', 75.99],
                   ['2017-01-01 03:00', '01/01/2017', 'Jan', 'Sun', '03:00', 60.76],
                   ['2017-01-01 04:00', '01/01/2017', 'Sep', 'Sat', '04:00', 49.01]],
                   columns=['timestamp', 'date', 'month', 'day', 'hour', 'price'])

def InvertKeyListDictionary(input_dict):
    return {w: k for k, v in input_dict.items() for w in v}

season_map = {'Spring': ['Mar', 'Apr', 'May'],
              'Summer': ['Jun', 'Jul', 'Aug'],
              'Autumn': ['Sep', 'Oct', 'Nov'],
              'Winter': ['Dec', 'Jan', 'Feb']}

weekend_map = {'Weekday': ['Mon', 'Tue', 'Wed', 'Thu', 'Fri'],
               'Weekend': ['Sat', 'Sun']}

month_map = InvertKeyListDictionary(season_map)
day_map = InvertKeyListDictionary(weekend_map)

df['season'] = df['month'].map(month_map)
df['daytype'] = df['day'].map(day_map)

df_groups = df.groupby(['season', 'daytype'])

df_groups.get_group(('Winter', 'Weekend'))

# output
# timestamp date month day hour price season daytype
# 2017-01-01 01:00 01/01/2017 Jan Sat 01:00 60.73 Winter Weekend 
# 2017-01-01 03:00 01/01/2017 Jan Sun 03:00 60.76 Winter Weekend 

【讨论】:

    【解决方案2】:

    如果想要按条件过滤数据,请使用 boolean indexing 和通过比较 dayofweek 和 isin 创建的布尔掩码来检查列表 L 中的成员资格:

    #changed timestamp values only for better sample
    print (df)
                timestamp        date month  day   hour  price
    0 2017-01-01 00:00:00  01/01/2017   Jan  Sun  00:00  60.23
    1 2017-01-03 00:00:00  01/01/2017   Jan  Sun  00:00  60.23
    2 2017-02-01 01:00:00  01/01/2017   Jan  Sun  01:00  60.73
    3 2017-02-05 01:00:00  01/01/2017   Jan  Sun  01:00  60.73
    4 2017-03-01 02:00:00  01/01/2017   Jan  Sun  02:00  75.99
    5 2017-04-01 03:00:00  01/01/2017   Jan  Sun  03:00  60.76
    6 2017-11-01 04:00:00  01/01/2017   Jan  Sun  04:00  49.01
    
    L = ['Nov','Dec','Jan','Feb']
    mask = (df['timestamp'].dt.dayofweek > 4) & (df['month'].isin(L))
    df1 = df[mask]
    print (df1)
                timestamp        date month  day   hour  price
    0 2017-01-01 00:00:00  01/01/2017   Jan  Sun  00:00  60.23
    3 2017-02-05 01:00:00  01/01/2017   Jan  Sun  01:00  60.73
    5 2017-04-01 03:00:00  01/01/2017   Jan  Sun  03:00  60.76
    

    如果需要 season 和日期类型的新列:

    df['season'] = (df['timestamp'].dt.month%12 + 3) // 3
    df['state'] = np.where(df['timestamp'].dt.dayofweek > 4, 'weekend','weekdays')
    print (df)
                timestamp        date month  day   hour  price  season     state
    0 2017-01-01 00:00:00  01/01/2017   Jan  Sun  00:00  60.23       1   weekend
    1 2017-01-03 00:00:00  01/01/2017   Jan  Sun  00:00  60.23       1  weekdays
    2 2017-02-01 01:00:00  01/01/2017   Jan  Sun  01:00  60.73       1  weekdays
    3 2017-02-05 01:00:00  01/01/2017   Jan  Sun  01:00  60.73       1   weekend
    4 2017-03-01 02:00:00  01/01/2017   Jan  Sun  02:00  75.99       2  weekdays
    5 2017-04-01 03:00:00  01/01/2017   Jan  Sun  03:00  60.76       2   weekend
    6 2017-11-01 04:00:00  01/01/2017   Jan  Sun  04:00  49.01       4  weekdays
    

    它可以用于 groupby 与聚合,例如sum:

    df2 = df.groupby(['season','state'], as_index=False)['price'].sum()
    print (df2)
       season     state   price
    0       1  weekdays  120.96
    1       1   weekend  120.96
    2       2  weekdays   75.99
    3       2   weekend   60.76
    4       4  weekdays   49.01
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2019-05-11
      • 2016-09-14
      • 2021-08-29
      • 2018-01-26
      • 2016-11-19
      • 2016-08-08
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多