【问题标题】:Select rows containing certain values from pandas dataframe从熊猫数据框中选择包含某些值的行
【发布时间】:2016-11-06 05:50:17
【问题描述】:

我有一个 pandas 数据框,其条目都是字符串:

   A     B      C
1 apple  banana pear
2 pear   pear   apple
3 banana pear   pear
4 apple  apple  pear

等等。我想选择包含某个字符串的所有行,比如'banana'。不知道每次会出现在哪一栏。当然,我可以编写一个 for 循环并遍历所有行。但是有没有更简单或更快的方法来做到这一点?

【问题讨论】:

  • 你也可以df[df.values == 'banana']
  • @JoeT.Boka,每场比赛都给我一行,所以如果一行有两个“香蕉”值,我会得到两行具有相同索引的行。不是不能处理的,但确实需要进一步处理。

标签: python pandas


【解决方案1】:

您可以通过将整个 df 与您的字符串进行比较来创建一个布尔掩码,并调用 dropna 传递参数 how='all' 以删除您的字符串未出现在所有列中的行:

In [59]:
df[df == 'banana'].dropna(how='all')

Out[59]:
        A       B    C
1     NaN  banana  NaN
3  banana     NaN  NaN

要测试多个值,您可以使用多个掩码:

In [90]:
banana = df[(df=='banana')].dropna(how='all')
banana

Out[90]:
        A       B    C
1     NaN  banana  NaN
3  banana     NaN  NaN

In [91]:    
apple = df[(df=='apple')].dropna(how='all')
apple

Out[91]:
       A      B      C
1  apple    NaN    NaN
2    NaN    NaN  apple
4  apple  apple    NaN

您可以使用index.intersection 仅对常见的索引值进行索引:

In [93]:
df.loc[apple.index.intersection(banana.index)]

Out[93]:
       A       B     C
1  apple  banana  pear

【讨论】:

  • 谢谢。如果我正在寻找一个字符串,这肯定有效。如果我想选择同时包含 'banana' 和 'apple' 的行怎么办?
  • 我不知道 pandas,但可能是这样的:df[df == 'banana', 'apple'].dropna(how='all')?
  • @Andromedae93 这给了我一个 TypeError
  • @mcglashan 我从未使用过 pandas,但 isin 函数应该可以工作。文档:pandas.pydata.org/pandas-docs/stable/generated/…
  • @JoeR 纯 numpy 方法总是更快,但 pandas 方法具有更好的类型和缺失数据处理,对于这个玩具示例,dtype 是同质的,那么纯 np 方法更胜一筹
【解决方案2】:

简介

在选择行的核心,我们需要一个一维掩码或一系列长度与df 相同的布尔元素,我们称之为mask。所以,最后使用df[mask],我们将在boolean-indexing 之后从df 中取出选定的行。

这是我们的起点df:

In [42]: df
Out[42]: 
        A       B      C
1   apple  banana   pear
2    pear    pear  apple
3  banana    pear   pear
4   apple   apple   pear

我。匹配一个字符串

现在,如果我们只需要匹配一个字符串,它是直接的元素相等:

In [42]: df == 'banana'
Out[42]: 
       A      B      C
1  False   True  False
2  False  False  False
3   True  False  False
4  False  False  False

如果我们需要在每一行中查找ANY 一个匹配项,请使用.any 方法:

In [43]: (df == 'banana').any(axis=1)
Out[43]: 
1     True
2    False
3     True
4    False
dtype: bool

选择相应的行:

In [44]: df[(df == 'banana').any(axis=1)]
Out[44]: 
        A       B     C
1   apple  banana  pear
3  banana    pear  pear

二。匹配多个字符串

1.搜索ANY匹配

这是我们的起点df:

In [42]: df
Out[42]: 
        A       B      C
1   apple  banana   pear
2    pear    pear  apple
3  banana    pear   pear
4   apple   apple   pear

NumPy 的 np.isin 可以在这里工作(或使用其他帖子中列出的 pandas.isin)从 df 的搜索字符串列表中获取所有匹配项。所以,假设我们在df 中寻找'pear' 或'apple':

In [51]: np.isin(df, ['pear','apple'])
Out[51]: 
array([[ True, False,  True],
       [ True,  True,  True],
       [False,  True,  True],
       [ True,  True,  True]])

# ANY match along each row
In [52]: np.isin(df, ['pear','apple']).any(axis=1)
Out[52]: array([ True,  True,  True,  True])

# Select corresponding rows with masking
In [56]: df[np.isin(df, ['pear','apple']).any(axis=1)]
Out[56]: 
        A       B      C
1   apple  banana   pear
2    pear    pear  apple
3  banana    pear   pear
4   apple   apple   pear

2。搜索ALL匹配

这是我们再次开始的df:

In [42]: df
Out[42]: 
        A       B      C
1   apple  banana   pear
2    pear    pear  apple
3  banana    pear   pear
4   apple   apple   pear

所以,现在我们正在寻找具有BOTH 的行,比如['pear','apple']。我们将使用NumPy-broadcasting:

In [66]: np.equal.outer(df.to_numpy(copy=False),  ['pear','apple']).any(axis=1)
Out[66]: 
array([[ True,  True],
       [ True,  True],
       [ True, False],
       [ True,  True]])

所以,我们有一个2 项目的搜索列表,因此我们有一个带有number of rows = len(df) 和number of cols = number of search items 的二维蒙版。因此,在上面的结果中,我们有'pear' 的第一个列和'apple' 的第二个列。

为了具体化,让我们为三个项目获取一个掩码['apple','banana', 'pear']:

In [62]: np.equal.outer(df.to_numpy(copy=False),  ['apple','banana', 'pear']).any(axis=1)
Out[62]: 
array([[ True,  True,  True],
       [ True, False,  True],
       [False,  True,  True],
       [ True, False,  True]])

此掩码的列分别为'apple','banana', 'pear'。

回到2搜索项目案例,我们之前有过:

In [66]: np.equal.outer(df.to_numpy(copy=False),  ['pear','apple']).any(axis=1)
Out[66]: 
array([[ True,  True],
       [ True,  True],
       [ True, False],
       [ True,  True]])

因为,我们正在每一行中寻找 ALL 匹配项:

In [67]: np.equal.outer(df.to_numpy(copy=False),  ['pear','apple']).any(axis=1).all(axis=1)
Out[67]: array([ True,  True, False,  True])

最后,选择行:

In [70]: df[np.equal.outer(df.to_numpy(copy=False),  ['pear','apple']).any(axis=1).all(axis=1)]
Out[70]: 
       A       B      C
1  apple  banana   pear
2   pear    pear  apple
4  apple   apple   pear

【讨论】:

  • 其实这个在搜索多个字符串的时候比较好用
【解决方案3】:

对于单个搜索值

df[df.values  == "banana"]

或

 df[df.isin(['banana'])]

对于多个搜索词:

  df[(df.values  == "banana")|(df.values  == "apple" ) ]

或

df[df.isin(['banana', "apple"])]

  #         A       B      C
  #  1   apple  banana    NaN
  #  2     NaN     NaN  apple
  #  3  banana     NaN    NaN
  #  4   apple   apple    NaN

来自 Divakar:返回包含两者的行。

select_rows(df,['apple','banana'])

 #         A       B     C
 #   0  apple  banana  pear

【讨论】:

  • 当我尝试时,最后一行实际上给了我一个空数据框
【解决方案4】:

如果您希望df 的所有行都包含values 中的任何 个值,请使用:

df[df.isin(values).any(1)]

例子:

In [2]: df                                                                                                                       
Out[2]: 
   0  1  2
0  7  4  9
1  8  2  7
2  1  9  7
3  3  8  5
4  5  1  1

In [3]: df[df.isin({1, 9, 123}).any(1)]                                                                                          
Out[3]: 
   0  1  2
0  7  4  9
2  1  9  7
4  5  1  1

【讨论】:

    猜你喜欢
    • 2018-09-19
    • 1970-01-01
    • 1970-01-01
    • 2019-10-03
    • 2020-01-12
    • 2018-08-05
    • 2022-01-21
    相关资源
    最近更新 更多