【问题标题】:Sorting grouped DataFrame column without changing index sorting在不更改索引排序的情况下对分组的 DataFrame 列进行排序
【发布时间】:2020-07-24 05:40:59
【问题描述】:
我有一个如下的df:
我只想要每年排名前 5 的国家/地区,但要保持年度递增。
首先我按年份和国家名称对 df 进行分组,然后运行以下代码:
df.sort_values(['year','hydro_total'], ascending=False).groupby(['year']).head(5)
结果并没有保持索引升序,相反,它也对年份索引进行了排序。如何获得前 5 名的国家并保持年度组的上升?
CSV 文件已上传 HERE 。
【问题讨论】:
标签:
python
pandas
numpy
dataframe
pandas-groupby
【解决方案1】:
您已经按year 和hydro_total 排序,两者均递减。您需要将年份排序为递增:
(df.sort_values(['year','hydro_total'],
ascending=[True,False])
.groupby('year').head(5)
)
输出:
country year hydro_total hydro_per_person
440 Japan 1971 7240000.0 0.06890
160 China 1971 2580000.0 0.00308
240 India 1971 2410000.0 0.00425
760 North Korea 1971 788000.0 0.05380
800 Pakistan 1971 316000.0 0.00518
... ... ... ... ...
199 China 2010 62100000.0 0.04630
279 India 2010 9840000.0 0.00803
479 Japan 2010 7070000.0 0.05590
1119 Turkey 2010 4450000.0 0.06120
839 Pakistan 2010 2740000.0 0.01580