【问题标题】:Using Groupby with Duplicated将 Groupby 与重复项一起使用
【发布时间】:2018-03-08 08:40:14
【问题描述】:

我对此进行了研究,但没有找到任何信息来执行以下操作。我需要

  1. 按一组列分组(此处为:studentid、主题、主题、课程)。
  2. 然后我需要在列的子集中找到重复的行 (这里:测试时间,响应时间)。
  3. 创建一个列来指示行(跨 2 列)是否重复

起始数据框

   studentid   subj   topic  lesson  testtime    responsetime
1  1           math   add    a       timestamp1  45sec
2  1           math   add    a       timestamp1  45sec
3  1           math   add    a       timestamp2  30sec
4  1           math   add    a       timestamp3  15sec
5  1           math   add    b       timestamp1  0sec
6  1           math   add    b       timestamp1  0sec
7  1           math   add    b       timestamp1  45sec
8  1           math   add    b       timestamp1  45sec

我尝试过的: 方法一: 使用在用户定义的函数中放置重复项来创建一个列,指示该行是否重复 - 错误:'function' object is not subscriptable

def check_dup(list):
    return df.duplicated([list],keep='first')

df_alt['dup_values'] = df.groupby(['studentidd', 'subj','topic','lesson']). apply(check_dup['testtime','responsetime'],axis=1)

方法二: 使用多索引,但问题是重复函数在索引行中查找重复项,而不是在单独的列集中('testtime','responsetime'):

  dfnew['dup_indicator'] = df.set_index(['studentidd', 'subj','topic','lesson']).
duplicated(['testtime','responsetime'],keep=False)

所需的数据帧

   studentid   subj   topic  lesson  testtime   responsetime dup_indicator
1  1           math   add    a       timestamp1  45sec             1
2  1           math   add    a       timestamp1  45sec             1
3  1           math   add    a       timestamp2  30sec             0
4  1           math   add    a       timestamp3  15sec             0
5  1           math   add    b       timestamp1  0sec              1 
6  1           math   add    b       timestamp1  0sec              1
7  1           math   add    b       timestamp1  45sec             1
8  1           math   add    b       timestamp1  45sec             1

【问题讨论】:

    标签: duplicates multiple-columns pandas-groupby multi-index


    【解决方案1】:

    您无需使用groupby 或修改索引即可完成您想做的事情。只需传入您要用于标识为重复项的所有列:

    > df
        studentid   subj    topic   lesson  testtime    responsetime
    1   1           math    add     a       timestamp1  45sec
    2   1           math    add     a       timestamp1  45sec
    3   1           math    add     a       timestamp2  30sec
    4   1           math    add     a       timestamp3  15sec
    5   1           math    add     b       timestamp1  0sec
    6   1           math    add     b       timestamp1  0sec
    7   1           math    add     b       timestamp1  45sec
    8   1           math    add     b       timestamp1  45sec
    
    > dup_cols = ['studentid', 'subj', 'topic', 'lesson', 'testtime', 'responsetime']
    > df.loc[df.duplicated(subset=dup_cols, keep=False), 'dup_indicator'] = 1
    > df['dup_indicator'].fillna(0, inplace=True)
    > df
    
        studentid   subj    topic   lesson  testtime    responsetime    dup_indicator
    1   1       math        add     a       timestamp1  45sec           1.0
    2   1       math        add     a       timestamp1  45sec           1.0
    3   1       math        add     a       timestamp2  30sec           0.0
    4   1       math        add     a       timestamp3  15sec           0.0
    5   1       math        add     b       timestamp1  0sec            1.0
    6   1       math        add     b       timestamp1  0sec            1.0
    7   1       math        add     b       timestamp1  45sec           1.0
    8   1       math        add     b       timestamp1  45sec           1.0
    

    分解步骤:

    • 查找df.duplicated 返回True 的所有行,在这种情况下,根据传递给subset 参数的所有行重复
    • 使用.loc过滤数据框以选择重复行
    • 创建一个新列 dup_indicator 并将 1 分配给重复的行
    • 使用fillna 将0 分配给非重复行

    【讨论】:

    • 这非常好用!这是一个比我想象的更简单的解决方案。
    • 一个后续问题 - 当我只想保留一个时,我添加 keep='first',但 value_counts 仍然为两个重复的行分配 df['dup_indicator] ==1。如何创建一个 df 只保留重复行的“第一个”。
    • 如果您不需要指标列,您可以使用以下内容:df.drop_duplicates(subset=dup_cols)。默认行为是保留重复行中的第一行
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-03-04
    • 2022-08-08
    • 2020-05-05
    • 2017-08-16
    • 2011-01-13
    • 2017-06-21
    相关资源
    最近更新 更多