【问题标题】:Filter rows containing certain values in the columns过滤列中包含特定值的行
【发布时间】:2023-03-29 21:38:01
【问题描述】:

我想通过数据框中的多列过滤掉包含特定值的行。

例如

    code tag number floor  note
1   1111  *   **     34     no 
2   2323  7   899     7     no
3   3677  #   900    11     no
4   9897  10  134    *      no
5    #    #   566    11     no
6   3677  55  908    11     no

我想过滤掉所有包含#, *, ** 列代码、标签、数字、楼层的行。

我想得到的是

    code tag number floor  note
1   1111  *   **     34     no 
3   3677  #   900    11     no
4   9897  10  134    *      no
5    #    #   566    11     no

我试图在数据框中使用 isin 方法,但它确实适用于一列,但不适用于多列。谢谢!

【问题讨论】:

    标签: python pandas dataframe filter


    【解决方案1】:

    选项 1
    假设没有其他预先存在的pir

    df[df.replace(['#', '*', '**'], 'pir').eq('pir').any(1)]
    
       code tag number floor note
    1  1111   *     **    34   no
    3  3677   #    900    11   no
    4  9897  10    134     *   no
    5     #   #    566    11   no
    

    选项 2
    令人讨厌的numpy 广播。一开始很快,但呈二次方缩放

    df[(df.values[None, :] == np.array(['*', '**', '#'])[:, None, None]).any(0).any(1)]
    
       code tag number floor note
    1  1111   *     **    34   no
    3  3677   #    900    11   no
    4  9897  10    134     *   no
    5     #   #    566    11   no
    

    选项 3
    不那么讨厌np.in1d

    df[np.in1d(df.values, ['*', '**', '#']).reshape(df.shape).any(1)]
    
       code tag number floor note
    1  1111   *     **    34   no
    3  3677   #    900    11   no
    4  9897  10    134     *   no
    5     #   #    566    11   no
    

    选项 4
    顶一下map

    df[list(
        map(bool,
            map({'*', '**', '#'}.intersection,
                map(set,
                    zip(*(df[c].values.tolist() for c in df)))))
    )]
    
       code tag number floor note
    1  1111   *     **    34   no
    3  3677   #    900    11   no
    4  9897  10    134     *   no
    5     #   #    566    11   no
    

    【讨论】:

      【解决方案2】:

      我认为您需要带有布尔索引的 apply、isin 和 any:

      list = ['#','*','**']
      cols = ['code','tag','number','floor']
      df[df[cols].apply(lambda x: x.isin(list).any(), axis=1)]
      

      输出:

         code tag number floor note
      1  1111   *     **    34   no
      3  3677   #    900    11   no
      4  9897  10    134     *   no
      5     #   #    566    11   no
      

      【讨论】:

        【解决方案3】:

        你也可以使用df.applymap:

        s = {'*', '**', '#'}
        df[df.applymap(lambda x: x in s).max(1)]
        
           code tag number floor note
        1  1111   *     **    34   no
        3  3677   #    900    11   no
        4  9897  10    134     *   no
        5     #   #    566    11   no
        

        piR suggested 一个疯狂的(但它有效!)替代方案:

        df[df.apply(set, 1) & {'*', '**', '#'}]
        
           code tag number floor note
        1  1111   *     **    34   no
        3  3677   #    900    11   no
        4  9897  10    134     *   no
        5     #   #    566    11   no
        

        【讨论】:

        • @piRSquared 令人震惊的是它确实有效。如果您对此感到满意,请将其添加为答案?
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2014-05-16
        • 2019-02-18
        • 1970-01-01
        • 1970-01-01
        • 2020-04-28
        相关资源
        最近更新 更多