【问题标题】:Pandas Groupby Head by Percentage of Row CountsPandas Groupby Head 按行数百分比
【发布时间】:2021-10-28 20:20:40
【问题描述】:

我有一个数据框:

state    city             score
CA       San Francisco    80
CA       San Francisco    90
...
NC       Raleigh          44
NY       New York City    22

我想做一个 groupby.head(),但不是整数值,而是选择每个州-城市组合的前 80%,按分数排序。 p>

因此,如果 CA, San Francisco 有 100 行,NC, Raleigh 有 20 行,则最终数据帧将包含 CA、San Francisco 的前 80 个分数行以及 NC、Raleigh 的前 16 个分数行。

所以最终结果代码可能类似于:

df.sort_values('score', ascending=False).groupby(['State', 'City']).head(80%)

谢谢!

【问题讨论】:

  • @It_is_Chris 这很接近,但它只返回随机选择的行的百分比,我需要按分数列对百分比进行排序。意识到我没有在我的原始帖子中澄清这一点,所以我现在将进行编辑

标签: python pandas pandas-groupby


【解决方案1】:
from io import StringIO
import pandas as pd

# sample data
s = """state,city,score
CA,San Francisco,80
CA,San Francisco,90
CA,San Francisco,30
CA,San Francisco,10
CA,San Francisco,70
CA,San Francisco,60
CA,San Francisco,50
CA,San Francisco,40
NC,Raleigh,44
NC,Raleigh,54
NC,Raleigh,64
NC,Raleigh,14
NY,New York City,22
NY,New York City,12
NY,New York City,32
NY,New York City,42
NY,New York City,52"""

df = pd.read_csv(StringIO(s))

sample = .8 # 80% 
# sort the values and create a groupby object
g = df.sort_values('score', ascending=False).groupby(['state', 'city']) 
# use list comprehension to iterate over each group
# for each group, calculate what 80% is
# in other words, the length of each group multiplied by .8
# you then use int to round down to the whole number
new_df = pd.concat([data.head(int(len(data)*sample)) for _,data in g])

   state           city  score
1     CA  San Francisco     90
0     CA  San Francisco     80
4     CA  San Francisco     70
5     CA  San Francisco     60
6     CA  San Francisco     50
7     CA  San Francisco     40
10    NC        Raleigh     64
9     NC        Raleigh     54
8     NC        Raleigh     44
16    NY  New York City     52
15    NY  New York City     42
14    NY  New York City     32
12    NY  New York City     22

【讨论】:

    【解决方案2】:

    使用nlargest并根据其长度计算每组的选定行数,即0.8 * len(group)

    res = (
        df.groupby(['State', 'City'], group_keys=False)
          .apply(lambda g: g.nlargest(int(0.8*len(g)), "Score"))
    )
    
    

    【讨论】:

      猜你喜欢
      • 2022-06-13
      • 1970-01-01
      • 2015-07-30
      • 2022-11-21
      • 2014-06-16
      • 2018-03-24
      • 1970-01-01
      相关资源
      最近更新 更多