【发布时间】:2020-09-01 14:54:51
【问题描述】:
我有以下熊猫系列:
Reducedset['% Renewable']
这给了我:
Asia China 19.7549
Japan 10.2328
India 14.9691
South Korea 2.27935
Iran 5.70772
North America United States 11.571
Canada 61.9454
Europe United Kingdom 10.6005
Russian Federation 17.2887
Germany 17.9015
France 17.0203
Italy 33.6672
Spain 37.9686
Australia Australia 11.8108
South America Brazil 69.648
Name: % Renewable, dtype: object
然后我将这个系列分为 5 个箱子:
binning = pd.cut(Top15['% Renewable'],5)
这给了我:
Asia China (15.753, 29.227]
Japan (2.212, 15.753]
India (2.212, 15.753]
South Korea (2.212, 15.753]
Iran (2.212, 15.753]
North America United States (2.212, 15.753]
Canada (56.174, 69.648]
Europe United Kingdom (2.212, 15.753]
Russian Federation (15.753, 29.227]
Germany (15.753, 29.227]
France (15.753, 29.227]
Italy (29.227, 42.701]
Spain (29.227, 42.701]
Australia Australia (2.212, 15.753]
South America Brazil (56.174, 69.648]
Name: % Renewable, dtype: category
Categories (5, interval[float64]): [(2.212, 15.753] < (15.753, 29.227] < (29.227, 42.701] <
(42.701, 56.174] < (56.174, 69.648]]
然后我对这些分箱数据进行分组,以计算每个分箱中的国家/地区数量:
Reduced = Reducedset.groupby(binning)['% Renewable'].agg(['count'])
这给了我:
% Renewable
(2.212, 15.753] 7
(15.753, 29.227] 4
(29.227, 42.701] 2
(42.701, 56.174] 0
(56.174, 69.648] 2
Name: count, dtype: int64
但是,索引已消失,我仍想保留“大陆”(外部索引)的索引。
因此,在 (% Renewable) 列的最左侧,它应该说:
Asia
North America
Europe
Australia
South America
当我尝试这样做时:
print(Reducedset['% Renewable'].groupby([Reducedset['% Renewable'].index.get_level_values(0),pd.cut(Reducedset['% Renewable'],5)]).count())
有效!
问题解决了!
【问题讨论】:
-
@Ben.T 实际上,我想要这个输出:count binning (2.212, 15.753] 7 (15.753, 29.227] 4 (29.227, 42.701] 2 (56.174, 69.648] 2 但是大陆包括索引
标签: python pandas dataframe indexing group-by