使用示例数据集可能更容易解释。
创建示例数据
假设我们有一列时间戳 date 和另一列我们要对其执行聚合的列 a。
df = pd.DataFrame({'date':pd.DatetimeIndex(['2012-1-1', '2012-6-1', '2015-1-1', '2015-2-1', '2015-3-1']),
'a':[9,5,1,2,3]}, columns=['date', 'a'])
df
date a
0 2012-01-01 9
1 2012-06-01 5
2 2015-01-01 1
3 2015-02-01 2
4 2015-03-01 3
有几种按年份分组的方法
- 将 dt 访问器与
year 属性结合使用
- 将
date放入索引并使用匿名函数访问年份
- 使用
resample方法
- 转换为 pandas 周期
.dt 具有 year 属性的访问器
当您有一列(而不是索引)熊猫时间戳时,您可以使用 dt 访问器访问更多额外的属性和方法。例如:
df['date'].dt.year
0 2012
1 2012
2 2015
3 2015
4 2015
Name: date, dtype: int64
我们可以使用它来形成我们的组并计算特定列上的一些聚合:
df.groupby(df['date'].dt.year)['a'].agg(['sum', 'mean', 'max'])
sum mean max
date
2012 14 7 9
2015 6 2 3
将日期放入索引并使用匿名函数访问年份
如果将日期列设置为索引,它将成为 DateTimeIndex,其属性和方法与 dt 访问器提供的普通列相同
df1 = df.set_index('date')
df1.index.year
Int64Index([2012, 2012, 2015, 2015, 2015], dtype='int64', name='date')
有趣的是,当使用 groupby 方法时,您可以向它传递一个函数。此函数将隐式传递 DataFrame 的索引。所以,我们可以从上面得到相同的结果:
df1.groupby(lambda x: x.year)['a'].agg(['sum', 'mean', 'max'])
sum mean max
2012 14 7 9
2015 6 2 3
使用resample 方法
如果您的日期列不在索引中,则必须使用on 参数指定该列。您还需要将offset alias 指定为字符串。
df.resample('AS', on='date')['a'].agg(['sum', 'mean', 'max'])
sum mean max
date
2012-01-01 14.0 7.0 9.0
2013-01-01 NaN NaN NaN
2014-01-01 NaN NaN NaN
2015-01-01 6.0 2.0 3.0
转换为 pandas 周期
您还可以将日期列转换为 pandas Period 对象。我们必须将偏移别名作为字符串传入,以确定 Period 的长度。
df['date'].dt.to_period('A')
0 2012
1 2012
2 2015
3 2015
4 2015
Name: date, dtype: object
然后我们可以将其用作一个组
df.groupby(df['date'].dt.to_period('Y'))['a'].agg(['sum', 'mean', 'max'])
sum mean max
2012 14 7 9
2015 6 2 3