【发布时间】:2018-03-08 08:40:14
【问题描述】:
我对此进行了研究,但没有找到任何信息来执行以下操作。我需要
- 按一组列分组(此处为:studentid、主题、主题、课程)。
- 然后我需要在列的子集中找到重复的行 (这里:测试时间,响应时间)。
- 创建一个列来指示行(跨 2 列)是否重复
起始数据框
studentid subj topic lesson testtime responsetime
1 1 math add a timestamp1 45sec
2 1 math add a timestamp1 45sec
3 1 math add a timestamp2 30sec
4 1 math add a timestamp3 15sec
5 1 math add b timestamp1 0sec
6 1 math add b timestamp1 0sec
7 1 math add b timestamp1 45sec
8 1 math add b timestamp1 45sec
我尝试过的: 方法一: 使用在用户定义的函数中放置重复项来创建一个列,指示该行是否重复 - 错误:'function' object is not subscriptable
def check_dup(list):
return df.duplicated([list],keep='first')
df_alt['dup_values'] = df.groupby(['studentidd', 'subj','topic','lesson']). apply(check_dup['testtime','responsetime'],axis=1)
方法二: 使用多索引,但问题是重复函数在索引行中查找重复项,而不是在单独的列集中('testtime','responsetime'):
dfnew['dup_indicator'] = df.set_index(['studentidd', 'subj','topic','lesson']).
duplicated(['testtime','responsetime'],keep=False)
所需的数据帧
studentid subj topic lesson testtime responsetime dup_indicator
1 1 math add a timestamp1 45sec 1
2 1 math add a timestamp1 45sec 1
3 1 math add a timestamp2 30sec 0
4 1 math add a timestamp3 15sec 0
5 1 math add b timestamp1 0sec 1
6 1 math add b timestamp1 0sec 1
7 1 math add b timestamp1 45sec 1
8 1 math add b timestamp1 45sec 1
【问题讨论】:
标签: duplicates multiple-columns pandas-groupby multi-index