【问题标题】:pandas how to compare similar rows and then drop by conditions熊猫如何比较相似的行然后按条件删除
【发布时间】:2021-08-17 22:08:43
【问题描述】:

我有一个数据框:

dfs = """
    contract  RB  BeginDate  ValIssueDate   EndDate   Valindex0
1  A00118  46   19000100      19880901  19841231          50
2  A00118  46   19850100      19880901  99999999          50
3  A00118  47   19000100      19880901  19831231          47
4  A00118  47   19840100      19880901  19841299          47
5  A00118  47   19850100      19880901  99999999          50
6  A00253  48   19000100      19820101  19811231          47
7  A00253  48   19820100      19820101  19841299          47
8  A00253  48   19850100      19820101  99999999          50
9  A00253  50   19000100      19820101  19781231          47
10 A00253  50   19790100      19820101  19841299          47
11 A00253  50   19850100      19820101  99999999          50
12 A00253  4L   20170101      19880901  99999999          39


"""
df = pd.read_csv(StringIO(dfs.strip()), sep='\s+', 
                  dtype={"RB": str, "BeginDate": int, "EndDate": int,'ValIssueDate':int,'Valindex0':int})

df:

   contract RB  BeginDate   ValIssueDate    EndDate Valindex0
1   A00118  46  19000100    19880901    19841231    50
2   A00118  46  19850100    19880901    99999999    50
3   A00118  47  19000100    19880901    19831231    47
4   A00118  47  19840100    19880901    19841299    47
5   A00118  47  19850100    19880901    99999999    50
6   A00253  48  19000100    19820101    19811231    47
7   A00253  48  19820100    19820101    19841299    47
8   A00253  48  19850100    19820101    99999999    50
9   A00253  50  19000100    19820101    19781231    47
10  A00253  50  19790100    19820101    19841299    47
11  A00253  50  19850100    19820101    99999999    50
12 A00253  4L   20170101      19880901  99999999    39

我想按此条件删除行:

如果该行与其他行具有相同的“合同”和“RB”,但其“ValIssueDate”不在 'BeginDate' 和 'EndDate',然后删除这一行。

注意最后一行,它有唯一的 RB,所以它不应该被删除。

index_names = df[ (df['ValIssueDate'] <= df['EndDate'] ) | (df['ValIssueDate'] >= df['BeginDate'])].index
# drop these given row
# indexes from dataFrame
df.drop(index_names, inplace = True)

这个方法只比较1行以内,但是如何根据我的条件比较不同的行呢?

输出应该是:

    contract  RB  BeginDate  ValIssueDate   EndDate   Valindex0
2  A00118  46   19850100      19880901  99999999          50
5  A00118  47   19850100      19880901  99999999          50
7  A00253  48   19820100      19820101  19841299          47
10 A00253  50   19790100      19820101  19841299          47
12 A00253  4L   20170101      19880901  99999999          39

【问题讨论】:

    标签: python pandas dataframe numpy


    【解决方案1】:

    不要删除不需要的行,而是保留需要的行。 您所做的布尔索引非常接近您的实际需要:

    1. 对于合约和 RB 重复的行,使用 ValIssueDate 上的条件
    2. 对于具有唯一合同和 RB 的行,保留所有行。
    df = df[((df.duplicated(subset = ["contract", "RB"], keep=False)) & 
             (df['ValIssueDate'] <= df['EndDate'] ) & 
             (df['ValIssueDate'] >= df['BeginDate'])) | 
            ~df.duplicated(subset = ["contract", "RB"], keep=False)]
    
    >>> df
       contract  RB  BeginDate  ValIssueDate   EndDate  Valindex0
    2    A00118  46   19850100      19880901  99999999         50
    5    A00118  47   19850100      19880901  99999999         50
    7    A00253  48   19820100      19820101  19841299         47
    10   A00253  50   19790100      19820101  19841299         47
    12   A00253  4L   20170101      19880901  99999999         39
    

    【讨论】:

    • 非常感谢您的回复,我刚刚更新了问题,请检查,我的错。
    • @William - 已编辑 :)
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-05-01
    • 2019-05-08
    • 1970-01-01
    • 1970-01-01
    • 2021-04-10
    相关资源
    最近更新 更多