【问题标题】:How can I keep all rows of a grouped pandas DataFrage meeting a certain criteria?如何使分组的熊猫 DataFrage 的所有行都满足特定条件?
【发布时间】:2019-11-13 04:05:51
【问题描述】:

对于带有组的 pandas DataFrame,我想保留所有行,直到第一次出现特定值(并丢弃所有其他行)。

MWE:

import pandas as pd
df = pd.DataFrame({'A' : ['foo', 'foo', 'foo', 'bar', 'bar', 'bar', 'tmp'],
                   'B' : [0, 1, 0, 0, 0, 1, 0],
                   'C' : [2.0, 5., 8., 1., 2., 9., 7.]})

给予

    A    B  C
0   foo  0  2.0
1   foo  1  5.0
2   foo  0  8.0
3   bar  0  1.0
4   bar  0  2.0
5   bar  1  9.0
6   tmp  0  7.0

我想保留每个组的所有行(A 是分组变量)直到B == 1(包括这一行)。所以,我想要的输出是

    A    B  C
0   foo  0  2.0
1   foo  1  5.0
3   bar  0  1.0
4   bar  0  2.0
5   bar  1  9.0
6   tmp  0  7.0

如何使分组 DataFrage 的所有行都满足特定条件?

我找到了how to drop specific groups not meeting a certain criteria (and keeping all other rows of all other groups),但没有找到如何删除所有组的特定行。我得到的最多的是获取每个组中行的索引,我想保留:

df.groupby('A').apply(lambda x: x['B'].cumsum().searchsorted(1))

导致

A
bar    2
foo    1
tmp    1

这还不够,因为它不返回实际数据(如果 tmp 的结果是 0 可能会更好)

【问题讨论】:

    标签: python-3.x search indexing pandas-groupby rows


    【解决方案1】:

    在阅读this question 关于groupby.apply 和groupby.aggregate 之间的区别之后,我意识到apply 适用于该组的所有列和行(因此是DataFrame?)。所以这是我应该应用于每个组的函数:

    def f(group):
        index = min(group['B'].cumsum().searchsorted(1), len(group))
        return group.iloc[0:index+1]
    

    通过运行df.groupby('A').apply(f),我得到了想要的结果:

                A       B   C
    A               
    bar     3   bar     0   1.0
            4   bar     0   2.0
            5   bar     1   9.0
    foo     0   foo     0   2.0
            1   foo     1   5.0
    tmp     6   tmp     0   7.0
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-08-22
      • 2020-09-28
      • 2022-07-22
      • 2019-11-15
      • 1970-01-01
      • 2012-01-07
      • 1970-01-01
      • 2020-10-11
      相关资源
      最近更新 更多