【问题标题】:Python - delete a row based on condition from a pandas.core.series.Series after groupbyPython - 在 groupby 之后根据条件从 pandas.core.series.Series 中删除一行
【发布时间】:2022-01-10 04:29:18
【问题描述】:

按两列case 和area 分组后,我有这个pandas.core.series.Series

case area
A 1 2494
2 2323
B 1 59243
2 27125
3 14

我只想保留 case A 中的区域,这意味着结果应该是这样的:

case area
A 1 2494
2 2323
B 1 59243
2 27125

我试过这段代码:

a = df['B'][~df['B'].index.isin(df['A'].index)].index
df['B'].drop(a)

它成功了,输出是:

但它并没有将它放在数据框中,它仍然是一样的。

当我分配 drop 的结果时,所有的值都变成了 NaN

df['B'] = df['B'].drop(a)

我该怎么办?

【问题讨论】:

  • 尝试添加.dropna()?
  • @mitoRibo 我不想删除案例 B 中的所有区域,我想删除案例 A 中不存在的区域
  • 感谢您的解释。我会通过删除你不想要的行然后分组来解决这个问题
  • @mitoRibo 分组后是不是不能去掉?

标签: python pandas group-by series drop


【解决方案1】:

分组后可以掉线,这是一种方法

import pandas
import numpy as np

np.random.seed(1)

ungroup_df = pd.DataFrame({
    'case':[
        'A','A','A','A','A','A',
        'A','A','A','A','A','A',
        'B','B','B','B','B','B',
        'B','B','B','B','B','B',
    ],
    'area':[
        1,2,1,2,1,2,
        1,2,1,2,1,2,
        1,2,3,1,2,3,
        1,2,3,1,2,3,
    ],
    'value': np.random.random(24),
})

df = ungroup_df.groupby(['case','area'])['value'].sum()
print(df)

#index into the multi-index to just the 'A' areas
#the ":" is saying any value at the first level (A or B)
#then the df.loc['A'].index is filtering to second level of index (area) that match A's
filt_df = df.loc[:,df.loc['A'].index]
print(filt_df)

测试df:

case  area
A     1       1.566114
      2       2.684593
B     1       1.983568
      2       1.806948
      3       2.079145
Name: value, dtype: float64

放下后的输出

case  area
A     1       1.566114
      2       2.684593
B     1       1.983568
      2       1.806948
Name: value, dtype: float64

【讨论】:

    猜你喜欢
    • 2023-03-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-02-23
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多