【问题标题】:compare columns in pivot table and add result比较数据透视表中的列并添加结果
【发布时间】:2018-08-24 13:31:04
【问题描述】:

我正在使用来自 http://senegal.opendataforafrica.org/SNVS2015/vital-statistics-of-senegal-2015 的关于塞内加尔人口的开放数据 csv。用 pandas 将其导入数据框(形状 17568,7)。

    region  regional-division   sex indicator                               Unit    Date    Value
0   Dakar   Total   Total       Populations (projection de 2008 à   2015)   Number  2008    2482294.0 
1   Dakar   Total   Total       Populations    (projection de 2008 à 2015)  Number  2009    2536959.0
2   Dakar   Total   Total       Populations (projection de 2008 à   2015)   Number  2010    2592191.0 
3   Dakar   Total   Total       Populations   (projection de 2008 à 2015)   Number  2011    2647751.0
4   Dakar   Total   Total       Populations (projection de 2008 à   2015)   Number  2012    2703203.0 
5   Dakar   Total   Total       Populations   (projection de 2008 à 2015)   Number  2013    2776787.0
6   Dakar   Total   Total       Populations (projection de 2008 à   2015)   Number  2014    2851556.0 
7   Dakar   Total   Total       Populations   (projection de 2008 à 2015)   Number  2015    2927422.0
8   Dakar   Total   Men         Populations (projection de 2008 à   2015)   Number  2008    1242463.0 
9   Dakar   Total   Men         Populations (projection   de 2008 à 2015)   Number  2009    1269764.0

然后做了

total_population_condition = (population['sex'] == 'Total') & (population['regional-division'] == 'Total')
total_population = population[total_population_condition]

除此之外

pivot_total_population = pd.pivot_table(total_population,values='Value',index=['region','sex'],columns='Date')

Pivot Table

现在的问题是:我想找出 2008 年至 2015 年间人口增长最快的 5 个地区。以及收缩率最高的 5 个地区。我试图使用“2008”值和“2015”值访问数据透视列,然后将后者划分为前者。然后将结果添加到数据框中。没能做到。我该怎么做?

更新:我刚刚想出了如何...

# compute growth first per region
pivot_total_population['growth'] = 
pivot_total_population.iloc[:,7]/pivot_total_population.iloc[:,0]

# then determine which are top 10 growing regions in terms of total population
pivot_total_population.sort_values(['growth'],ascending=False).head(10)

# then determine which are top 10 shrinking regions in terms of total population
pivot_total_population.sort_values(['growth'],ascending=True).head(10)

【问题讨论】:

  • 您应该考虑将您对问题的解决方案转化为答案的机会(是的,很高兴您回答自己的问题!)。接下来你可以批准你自己的答案,所以强调它是你问题的解决方案——此外,它可以得到其他用户的赞同,独立于你的赞同”会收到你的问题。

标签: python dataframe pivot-table


【解决方案1】:

找到了答案(感谢 gboffi 给新手的流程提示;-))

# compute growth first per region
pivot_total_population['growth'] = 
pivot_total_population.iloc[:,7]/pivot_total_population.iloc[:,0]

# then determine which are top 10 growing regions in terms of total population
pivot_total_population.sort_values(['growth'],ascending=False).head(10)

# then determine which are top 10 shrinking regions in terms of total population
pivot_total_population.sort_values(['growth'],ascending=True).head(10)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2020-06-26
    • 2015-09-26
    • 1970-01-01
    • 1970-01-01
    • 2016-04-29
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多