【问题标题】:Sorting the Pandas DataFrame Describe output对 Pandas 数据帧描述输出进行排序
【发布时间】:2019-09-23 19:51:34
【问题描述】:

我正在尝试使用计数对describe() 的输出进行排序。不知道怎么解决。

尝试了sort_by 和.loc,但它们都不能用于对描述的输出进行排序。

需要编辑下面这行代码:

df.groupby("Disease_Category")['Approved_Amt'].describe().reset_index()

电流输出

Disease_Category count mean std min 25% 50% 75% max
1 Disease1 5.0 82477.600000 51744.487632 30461.0 58318.00 72201.0 83408.00 168000.0 

2 Disease2 190.0 35357.163158 46268.552683 1707.0 13996.25 22186.0 36281.75 331835.0 

期望的输出

Disease_Category count mean std min 25% 50% 75% max
1 Disease2 190.0 35357.163158 46268.552683 1707.0 13996.25 22186.0 36281.75 331835.0 

2 Disease1 5.0 82477.600000 51744.487632 30461.0 58318.00 72201.0 83408.00 168000.0

【问题讨论】:

  • 排序函数为sort()和sort_values()。你从哪里得到sort_by()?

标签: python-3.x


【解决方案1】:

这应该可行。

import pandas as pd

df = pd.DataFrame({'Disease' : ['Disease1', 'Disease2'],
                   'Category' : [5,190],
                   'count' : [82477, 35357],
                   'mean' : [51744, 46268],
                   'std' : [30461, 1707],
                   'etc' : [1,2]})

print(df)
#   Category   Disease  count  etc   mean    std
#0         5  Disease1  82477    1  51744  30461
#1       190  Disease2  35357    2  46268   1707

# invert rows of dataframe so last row becomes the first
df = df.reindex(index = df.index[::-1])

df = df.reset_index()

#   index  Category   Disease  count  etc   mean    std
#0      1       190  Disease2  35357    2  46268   1707
#1      0         5  Disease1  82477    1  51744  30461

【讨论】:

  • 如果您正在反转行序列,那么它不会排序。动作背后没有任何语义来确定行的顺序。如果它有效,那基本上是因为行首先已经排序并且通过反转它,您将它从升序更改为降序等。
  • 同意,反转索引是一个理想的解决方案,因为我正在查看排序“计数”并且我有大约 1500 行要排序。
【解决方案2】:

使用sort_values()。 Documentation.

import pandas as pd

df1 = df.groupby("Disease_Category")['Approved_Amt'].describe().reset_index()
>>>df1
  Disease_Category  count   mean    std    min       25%    50%       75%     max
0         Disease1      5  82477  51744  30461  58318.00  72201  83408.00  168000
1         Disease2    190  35357  46268   1707  13996.25  22186  36281.75  331835

>>>df1.sort_values('count', ascending=False)
  Disease_Category  count   mean    std    min       25%    50%       75%     max
1         Disease2    190  35357  46268   1707  13996.25  22186  36281.75  331835
0         Disease1      5  82477  51744  30461  58318.00  72201  83408.00  168000

【讨论】:

  • 谢谢。有没有办法在与 describe 命令相同的行中添加 sort_values。
  • @Shins 只要结构仍然是数据框,您就应该能够将其链接起来。 df.groupby("Disease_Category")['Approved_Amt'].describe().reset_index().sort_values('count', ascending=False)。你只需要权衡它的可读性等。
  • 您好,我遇到了这个错误:# 检查重复项 KeyError: 'count',知道吗?
  • @harmoniuscool 你是否也想描述一个名为“count”的字段?
猜你喜欢
  • 2017-06-30
  • 2017-03-08
  • 1970-01-01
  • 2019-04-26
  • 2015-09-16
  • 2013-10-15
  • 2017-08-21
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多