【发布时间】:2018-06-29 08:21:25
【问题描述】:
这是Pandas: How to subset (and sum) top N observations within subcategories? 的后续问题,其中演示了如何在此数据框中找到每年前 3 个月的总和:
示例数据框
year month passengers
0 1949 January 112
1 1949 February 118
2 1949 March 132
3 1949 April 129
4 1949 May 121
5 1949 June 135
.
.
.
137 1960 June 535
138 1960 July 622
139 1960 August 606
140 1960 September 508
141 1960 October 461
142 1960 November 390
143 1960 December 432
所以你最终会得到这个:
year passengers
0 1949 432
1 1950 498
2 1951 582
3 1952 690
4 1953 779
5 1954 859
6 1955 1026
7 1956 1192
8 1957 1354
9 1958 1431
10 1959 1579
11 1960 176
数字432 for 1949是148+148+136 for the months July, August and September.的总和
我现在的问题是:
是否可以进行相同的计算,同时将相应的子类别作为列表保留在自己的列中?
期望的输出
(我只检查了 1949 年的实际总和。1950 年是弥补的):
year passengers months
0 1949 432 July, August, September
1 1950 498 August, September, December
2 1951 582 .
3 1952 690 .
4 1953 779 .
5 1954 859 .
6 1955 1026 .
7 1956 1192 .
8 1957 1354 .
9 1958 1431 .
10 1959 1579 .
11 1960 176 .
可重现的代码和数据:
import pandas as pd
import seaborn as sns
df = sns.load_dataset('flights')
print(df.head())
df2 = df.groupby('year')['passengers'].apply(lambda x: x.nlargest(3).sum()).reset_index()
print(df2.head())
df:
year month passengers
0 1949 January 112
1 1949 February 118
2 1949 March 132
3 1949 April 129
4 1949 May 121
df2:
year passengers
0 1949 432
1 1950 498
2 1951 582
3 1952 690
4 1953 779
感谢您的任何建议!
【问题讨论】: