简介
在选择行的核心,我们需要一个一维掩码或一系列长度与df 相同的布尔元素,我们称之为mask。所以,最后使用df[mask],我们将在boolean-indexing 之后从df 中取出选定的行。
这是我们的起点df:
In [42]: df
Out[42]:
A B C
1 apple banana pear
2 pear pear apple
3 banana pear pear
4 apple apple pear
我。匹配一个字符串
现在,如果我们只需要匹配一个字符串,它是直接的元素相等:
In [42]: df == 'banana'
Out[42]:
A B C
1 False True False
2 False False False
3 True False False
4 False False False
如果我们需要在每一行中查找ANY 一个匹配项,请使用.any 方法:
In [43]: (df == 'banana').any(axis=1)
Out[43]:
1 True
2 False
3 True
4 False
dtype: bool
选择相应的行:
In [44]: df[(df == 'banana').any(axis=1)]
Out[44]:
A B C
1 apple banana pear
3 banana pear pear
二。匹配多个字符串
1.搜索ANY匹配
这是我们的起点df:
In [42]: df
Out[42]:
A B C
1 apple banana pear
2 pear pear apple
3 banana pear pear
4 apple apple pear
NumPy 的 np.isin 可以在这里工作(或使用其他帖子中列出的 pandas.isin)从 df 的搜索字符串列表中获取所有匹配项。所以,假设我们在df 中寻找'pear' 或'apple':
In [51]: np.isin(df, ['pear','apple'])
Out[51]:
array([[ True, False, True],
[ True, True, True],
[False, True, True],
[ True, True, True]])
# ANY match along each row
In [52]: np.isin(df, ['pear','apple']).any(axis=1)
Out[52]: array([ True, True, True, True])
# Select corresponding rows with masking
In [56]: df[np.isin(df, ['pear','apple']).any(axis=1)]
Out[56]:
A B C
1 apple banana pear
2 pear pear apple
3 banana pear pear
4 apple apple pear
2。搜索ALL匹配
这是我们再次开始的df:
In [42]: df
Out[42]:
A B C
1 apple banana pear
2 pear pear apple
3 banana pear pear
4 apple apple pear
所以,现在我们正在寻找具有BOTH 的行,比如['pear','apple']。我们将使用NumPy-broadcasting:
In [66]: np.equal.outer(df.to_numpy(copy=False), ['pear','apple']).any(axis=1)
Out[66]:
array([[ True, True],
[ True, True],
[ True, False],
[ True, True]])
所以,我们有一个2 项目的搜索列表,因此我们有一个带有number of rows = len(df) 和number of cols = number of search items 的二维蒙版。因此,在上面的结果中,我们有'pear' 的第一个列和'apple' 的第二个列。
为了具体化,让我们为三个项目获取一个掩码['apple','banana', 'pear']:
In [62]: np.equal.outer(df.to_numpy(copy=False), ['apple','banana', 'pear']).any(axis=1)
Out[62]:
array([[ True, True, True],
[ True, False, True],
[False, True, True],
[ True, False, True]])
此掩码的列分别为'apple','banana', 'pear'。
回到2搜索项目案例,我们之前有过:
In [66]: np.equal.outer(df.to_numpy(copy=False), ['pear','apple']).any(axis=1)
Out[66]:
array([[ True, True],
[ True, True],
[ True, False],
[ True, True]])
因为,我们正在每一行中寻找 ALL 匹配项:
In [67]: np.equal.outer(df.to_numpy(copy=False), ['pear','apple']).any(axis=1).all(axis=1)
Out[67]: array([ True, True, False, True])
最后,选择行:
In [70]: df[np.equal.outer(df.to_numpy(copy=False), ['pear','apple']).any(axis=1).all(axis=1)]
Out[70]:
A B C
1 apple banana pear
2 pear pear apple
4 apple apple pear