【问题标题】:pandas finding duplicate rows with different label熊猫查找具有不同标签的重复行
【发布时间】:2022-10-31 11:17:22
【问题描述】:

我有一个我想对标记数据进行完整性检查的情况。我有数百个特征,想找到具有相同特征但标签不同的点。然后应该对这些发现的不一致标签簇进行编号并放入新的数据框中。 这并不难,但我想知道最优雅的解决方案是什么。 这里有一个例子:

import pandas as pd

df = pd.DataFrame({
    "feature_1" : [0,0,0,4,4,2],
    "feature_2" : [0,5,5,1,1,3],
    "label" : ["A","A","B","B","D","A"]
})

result_df = pd.DataFrame({
    "cluster_index" : [0,0,1,1],
    "feature_1" : [0,0,4,4],
    "feature_2" : [5,5,1,1],
    "label" : ["A","B","B","D"]
})

【问题讨论】:

    标签: pandas


    【解决方案1】:

    为了获得您想要的输出(重复数据删除和 cluster_index),您可以使用groupby 方法:

    g = df.groupby(['feature_1', 'feature_2'])['label']
    
    (df.assign(cluster_index=g.ngroup()) # get group name
       .loc[g.transform('size').gt(1)]   # filter the non-duplicates
       # line below only to have a nice cluster_index range (0,1…)
       .assign(cluster_index= lambda d: d['cluster_index'].factorize()[0])
    )
    

    输出:

       feature_1  feature_2 label  cluster_index
    1          0          5     A              0
    2          0          5     B              0
    3          4          1     B              1
    4          4          1     D              1
    

    【讨论】:

      【解决方案2】:

      首先获取每个feature 列的所有重复值,然后在必要时删除所有列的重复值(此处在示例数据中不需要),最后为组索引添加GroupBy.ngroup:

      df = df[df.duplicated(['feature_1','feature_2'],keep=False)].drop_duplicates()
      
      df['cluster_index'] = df.groupby(['feature_1', 'feature_2'])['label'].ngroup()
      print (df)
         feature_1  feature_2 label  cluster_index
      1          0          5     A              0
      2          0          5     B              0
      3          4          1     B              1
      4          4          1     D              1
      

      【讨论】:

        【解决方案3】:
        df1.assign(col1=df1.duplicated(subset='feature_1,feature_2'.split(','),keep=False))
            .assign(col2=df1.duplicated(subset='feature_1,feature_2,label'.split(','),keep=False))
            .loc[lambda dd:dd.col1&~dd.col2]
        
          feature_1  feature_2 label  col1   col2
        1          0          5     A  True  False
        2          0          5     B  True  False
        3          4          1     B  True  False
        4          4          1     D  True  False
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2017-01-31
          • 1970-01-01
          • 2017-03-26
          • 1970-01-01
          • 1970-01-01
          • 2018-04-21
          • 2021-12-30
          • 2020-08-21
          相关资源
          最近更新 更多