【问题标题】:String contains across two pandas series字符串包含两个熊猫系列
【发布时间】:2018-08-08 23:17:44
【问题描述】:

我在 pandas 数据框中有一个包含一些字符串的系列。我想在相邻列中搜索该字符串是否存在。

在下面的示例中,我想搜索“choice”系列中的字符串是否包含在“fruit”系列中,在新列“choice_match”中返回 true (1) 或 false (0)。

示例数据框:

import pandas as pd
d = {'ID': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 'fruit': [
'apple, banana', 'apple', 'apple', 'pineapple', 'apple, pineapple',            'orange', 'apple, orange', 'orange', 'banana', 'apple, peach'],
'choice': ['orange', 'orange', 'apple', 'pineapple', 'apple', 'orange',  'orange', 'orange', 'banana', 'banana']}
df = pd.DataFrame(data=d)

所需的数据帧:

import pandas as pd
d = {'ID': [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], 'fruit': [
'apple, banana', 'apple', 'apple', 'pineapple', 'apple, pineapple',   'orange', 'apple, orange', 'orange', 'banana', 'apple, peach'],
'choice': ['orange', 'orange', 'apple', 'pineapple', 'apple', 'orange',      'orange', 'orange', 'banana', 'banana'],
'choice_match': [0, 0, 1, 1, 1, 1, 1, 1, 1, 0]}
df = pd.DataFrame(data=d)

【问题讨论】:

    标签: python string pandas dataframe


    【解决方案1】:
    In [75]: df['choice_match'] = (df['fruit']
                                     .str.split(',\s*', expand=True)
                                     .eq(df['choice'], axis=0)
                                     .any(1).astype(np.int8))
    
    In [76]: df
    Out[76]:
       ID     choice             fruit  choice_match
    0   1     orange     apple, banana             0
    1   2     orange             apple             0
    2   3      apple             apple             1
    3   4  pineapple         pineapple             1
    4   5      apple  apple, pineapple             1
    5   6     orange            orange             1
    6   7     orange     apple, orange             1
    7   8     orange            orange             1
    8   9     banana            banana             1
    9  10     banana      apple, peach             0
    

    一步一步:

    In [78]: df['fruit'].str.split(',\s*', expand=True)
    Out[78]:
               0          1
    0      apple     banana
    1      apple       None
    2      apple       None
    3  pineapple       None
    4      apple  pineapple
    5     orange       None
    6      apple     orange
    7     orange       None
    8     banana       None
    9      apple      peach
    
    In [79]: df['fruit'].str.split(',\s*', expand=True).eq(df['choice'], axis=0)
    Out[79]:
           0      1
    0  False  False
    1  False  False
    2   True  False
    3   True  False
    4   True  False
    5   True  False
    6  False   True
    7   True  False
    8   True  False
    9  False  False
    
    In [80]: df['fruit'].str.split(',\s*', expand=True).eq(df['choice'], axis=0).any(1)
    Out[80]:
    0    False
    1    False
    2     True
    3     True
    4     True
    5     True
    6     True
    7     True
    8     True
    9    False
    dtype: bool
    
    In [81]: df['fruit'].str.split(',\s*', expand=True).eq(df['choice'], axis=0).any(1).astype(np.int8)
    Out[81]:
    0    0
    1    0
    2    1
    3    1
    4    1
    5    1
    6    1
    7    1
    8    1
    9    0
    dtype: int8
    

    【讨论】:

      【解决方案2】:

      这是一种方法:

      df['choice_match'] = df.apply(lambda row: row['choice'] in row['fruit'].split(','),\
                                    axis=1).astype(int)
      

      说明

      • df.apply 和 axis=1 循环遍历每一行并应用逻辑;它接受匿名的lambda 函数。
      • row['fruit'].split(',') 从fruit 列创建一个列表。这是必要的,例如,apple 不会在 pineapple 中考虑。
      • astype(int) 是将布尔值转换为整数以进行显示所必需的。

      【讨论】:

      • 谢谢你的解释,非常有帮助:)
      【解决方案3】:

      选项 1
      使用 Numpy 的 find
      当find没有找到值时,返回-1

      from numpy.core.defchararray import find
      
      choice = df.choice.values.astype(str)
      fruit = df.fruit.values.astype(str)
      
      df.assign(choice_match=(find(fruit, choice) > -1).astype(np.uint))
      
         ID     choice             fruit  choice_match
      0   1     orange     apple, banana             0
      1   2     orange             apple             0
      2   3      apple             apple             1
      3   4  pineapple         pineapple             1
      4   5      apple  apple, pineapple             1
      5   6     orange            orange             1
      6   7     orange     apple, orange             1
      7   8     orange            orange             1
      8   9     banana            banana             1
      9  10     banana      apple, peach             0
      

      选项 2
      设置逻辑
      sets < 是严格子集,<= 是子集。让自己成为sets 中的pd.Series 并使用<= 来确定一列的集合是否是另一列集合的子集。

      choice = df.choice.apply(lambda x: set([x]))
      fruit = df.fruit.str.split(', ').apply(set)
      
      df.assign(choice_match=(choice <= fruit).astype(np.uint))
      
         ID     choice             fruit  choice_match
      0   1     orange     apple, banana             0
      1   2     orange             apple             0
      2   3      apple             apple             1
      3   4  pineapple         pineapple             1
      4   5      apple  apple, pineapple             1
      5   6     orange            orange             1
      6   7     orange     apple, orange             1
      7   8     orange            orange             1
      8   9     banana            banana             1
      9  10     banana      apple, peach             0
      

      选项 3
      灵感来自@Wen's answer
      使用get_dummies 和max

      c = pd.get_dummies(df.choice)
      f = df.fruit.str.get_dummies(', ')
      df.assign(choice_match=pd.DataFrame.mul(*c.align(f, 'inner')).max(1))
      
         ID     choice             fruit  choice_match
      0   1     orange     apple, banana             0
      1   2     orange             apple             0
      2   3      apple             apple             1
      3   4  pineapple         pineapple             1
      4   5      apple  apple, pineapple             1
      5   6     orange            orange             1
      6   7     orange     apple, orange             1
      7   8     orange            orange             1
      8   9     banana            banana             1
      9  10     banana      apple, peach             0
      

      【讨论】:

        【解决方案4】:

        嗯,找个有趣的方法get_dummies

        (df.fruit.str.replace(' ','').str.get_dummies(',')+df.choice.str.get_dummies()).gt(1).any(1)
        Out[726]: 
        0    False
        1    False
        2     True
        3     True
        4     True
        5     True
        6     True
        7     True
        8     True
        9    False
        dtype: bool
        

        分配回来后

        df['New']=(df.fruit.str.replace(' ','').str.get_dummies(',')+df.choice.str.get_dummies()).gt(1).any(1).astype(int)
        df
        Out[728]: 
           ID     choice             fruit  New
        0   1     orange     apple, banana    0
        1   2     orange             apple    0
        2   3      apple             apple    1
        3   4  pineapple         pineapple    1
        4   5      apple  apple, pineapple    1
        5   6     orange            orange    1
        6   7     orange     apple, orange    1
        7   8     orange            orange    1
        8   9     banana            banana    1
        9  10     banana      apple, peach    0
        

        【讨论】:

        • 我决定在我的答案中添加一个受此启发的选项。很好的答案!
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2012-08-05
        相关资源
        最近更新 更多