【问题标题】:Efficient grouping in pandas based on another Series基于另一个系列的熊猫高效分组
【发布时间】:2017-03-21 11:53:22
【问题描述】:

我需要执行一个基于DataFrame 中另一个布尔列的分组操作。在示例中最容易看到:我有以下DataFrame:

    b          id   
0   False      0
1   True       0
2   False      0
3   False      1
4   True       1
5   True       2
6   True       2
7   False      3
8   True       4
9   True       4
10  False      4

并且想要获得一个列,如果 b 列是 True 并且它是给定 id 的最后一次为 True,则其元素为 True:

    b          id    lastMention
0   False      0     False
1   True       0     True
2   False      0     False
3   False      1     False
4   True       1     False
5   True       2     True
6   True       3     True
7   False      3     False
8   True       4     False
9   True       4     True
10  False      4     False

我有一个代码可以实现这一点,虽然效率低:

def lastMentionFun(df):
    b = df['b']
    a = b.sum()
    if a > 0:
        maxInd = b[b].index.max()
        df.loc[maxInd, 'lastMention'] = True
    return df

df['lastMention'] = False
df = df.groupby('id').apply(lastMentionFun)

有人可以提出正确的pythonic方法来快速完成这项工作吗?

【问题讨论】:

    标签: python python-3.x pandas dataframe group-by


    【解决方案1】:

    您可以先过滤列b 中为True 的值,然后使用groupby 和聚合max 获得max 索引值:

    print (df[df.b].reset_index().groupby('id')['index'].max())
    id
    0    1
    1    4
    2    6
    4    9
    Name: index, dtype: int64
    

    然后用loc的索引值替换值False:

    df['lastMention'] = False
    df.loc[df[df.b].reset_index().groupby('id')['index'].max(), 'lastMention'] = True
    
    print (df)
            b  id  lastMention
    0   False   0        False
    1    True   0         True
    2   False   0        False
    3   False   1        False
    4    True   1         True
    5    True   2        False
    6    True   2         True
    7   False   3        False
    8    True   4        False
    9    True   4         True
    10  False   4        False
    

    另一种解决方案 - 使用 groupby 和 apply 获取 max 索引值,然后使用 isin 测试索引中值的成员资格 - 输出为 boolean Series:

    print (df[df.b].groupby('id').apply(lambda x: x.index.max()))
    id
    0    1
    1    4
    2    6
    4    9
    dtype: int64
    
    df['lastMention'] = df.index.isin(df[df.b].groupby('id').apply(lambda x: x.index.max()))
    print (df)
            b  id lastMention
    0   False   0       False
    1    True   0        True
    2   False   0       False
    3   False   1       False
    4    True   1        True
    5    True   2       False
    6    True   2        True
    7   False   3       False
    8    True   4       False
    9    True   4        True
    10  False   4       False
    

    【讨论】:

      【解决方案2】:

      不确定这是否是最有效的方法,但它只使用内置函数(主要是“cumsum”,然后是 max 来检查它是否等于最后一个 - pd.merge 只是用来放置max 回到表中,也许有更好的方法来做到这一点?)。

      df['cum_b']=df.groupby('id', as_index=False).cumsum()
      df = pd.merge(df, df[['id','cum_b']].groupby('id', as_index=False).max(), how='left', on='id', suffixes=('','_max'))
      df['lastMention'] = np.logical_and(df.b, df.cum_b == df.cum_b_max)
      

      附:您在示例中指定的数据框从第一个到第二个 sn-p 略有变化,希望我正确解释了您的请求!

      【讨论】:

        猜你喜欢
        • 2017-08-12
        • 2022-08-13
        • 2013-08-01
        • 2021-09-10
        • 2018-02-18
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2023-01-20
        相关资源
        最近更新 更多