【发布时间】:2019-11-13 04:05:51
【问题描述】:
对于带有组的 pandas DataFrame,我想保留所有行,直到第一次出现特定值(并丢弃所有其他行)。
MWE:
import pandas as pd
df = pd.DataFrame({'A' : ['foo', 'foo', 'foo', 'bar', 'bar', 'bar', 'tmp'],
'B' : [0, 1, 0, 0, 0, 1, 0],
'C' : [2.0, 5., 8., 1., 2., 9., 7.]})
给予
A B C
0 foo 0 2.0
1 foo 1 5.0
2 foo 0 8.0
3 bar 0 1.0
4 bar 0 2.0
5 bar 1 9.0
6 tmp 0 7.0
我想保留每个组的所有行(A 是分组变量)直到B == 1(包括这一行)。所以,我想要的输出是
A B C
0 foo 0 2.0
1 foo 1 5.0
3 bar 0 1.0
4 bar 0 2.0
5 bar 1 9.0
6 tmp 0 7.0
如何使分组 DataFrage 的所有行都满足特定条件?
我找到了how to drop specific groups not meeting a certain criteria (and keeping all other rows of all other groups),但没有找到如何删除所有组的特定行。我得到的最多的是获取每个组中行的索引,我想保留:
df.groupby('A').apply(lambda x: x['B'].cumsum().searchsorted(1))
导致
A
bar 2
foo 1
tmp 1
这还不够,因为它不返回实际数据(如果 tmp 的结果是 0 可能会更好)
【问题讨论】:
标签: python-3.x search indexing pandas-groupby rows