【问题标题】:Python - Searching a string within a dataframe from a listPython - 从列表中搜索数据框中的字符串
【发布时间】:2019-03-06 17:00:53
【问题描述】:

我有以下清单:

search_list = ['STEEL','IRON','GOLD','SILVER']

我需要在数据框 (df) 中搜索:

      a    b             
0    123   'Blah Blah Steel'
1    456   'Blah Blah Blah'
2    789   'Blah Blah Gold'

并将匹配的行插入到一个新的数据框(newdf)中,从列表中添加一个带有匹配词的新列:

      a    b                   c
0    123   'Blah Blah Steel'   'STEEL'
1    789   'Blah Blah Gold'    'GOLD'

我可以使用下面的代码来提取匹配的行:

newdf=df[df['b'].str.upper().str.contains('|'.join(search_list),na=False)]

但我不知道如何将列表中的匹配词添加到 c 列中。

我认为匹配需要以某种方式捕获列表中匹配单词的索引,然后使用索引号提取值,但我不知道该怎么做。

任何帮助或指点将不胜感激

谢谢

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    你也可以这样做:

    import pandas as pd
    
    search_list = ('STEEL','IRON','GOLD','SILVER')
    
    df = pd.DataFrame({'a':[123,456,789],'b':['blah blah Steel','blah blah blah','blah blah Gold']})
    
    df.assign(c = df['b'].apply(lambda x: [j for j in x.split() if j.upper() in search_list]))
    

    更新速度

    
    import pandas as pd
    
    search_list = set(['STEEL','IRON','GOLD','SILVER'])
    
    df = pd.DataFrame({'a':[123,456,789],'b':['blah blah Steel','blah blah blah','blah blah Gold']})
    
    df.assign(c = lambda d: d['b'].str.upper().str.split().map(lambda x: set(x).intersection(search_list)))
    

    结果:

    【讨论】:

    • 使用这个答案,我得到一个错误“'float' object has no attribute 'split'”-知道为什么会这样吗?
    • 用str(x).split替换x.split
    • 当我进行替换时,我现在收到错误 TypeError: 'builtin_function_or_method' object is not iterable。
    • 你的列数据是什么数据类型,是否有缺失值?
    • 导入时会自动填充一个空白值。我应该用其他东西填充它吗?当我执行 df.dtypes 时,“b”列显示“对象”。对该列进行的唯一操作是导入时的 str.upper()。
    【解决方案2】:

    您可以使用set.intersection 查找b 列中出现的单词:

    search_list = set(['STEEL','IRON','GOLD','SILVER'])
    df['c'] = df['b'].apply(lambda x: set.intersection(set(x.upper().split(' ')), search_list))
    

    输出:

         a                b        c
    0  123  Blah Blah Steel  {STEEL}
    1  456   Blah Blah Blah       {}
    2  789   Blah Blah Gold   {GOLD}
    

    如果您想删除没有匹配的行,请使用df[df['c'].astype(bool)]

         a                b        c
    0  123  Blah Blah Steel  {STEEL}
    2  789   Blah Blah Gold   {GOLD}
    

    【讨论】:

      【解决方案3】:

      一种方法是

      def get_word(my_string):
          for word in search_list:
               if word.lower() in my_string.lower():
                     return word
          return None
      
      new_df["c"]= new_df["b"].apply(get_word)
      

      你也可以做一些类似的事情

      new_df["c"]= new_df["b"].apply(lambda my_string: [word for word in search_list if word.lower() in my_string.lower()][0])
      

      对于第一个,您可以选择先将c 列添加到df,然后过滤掉Nones,而第二个将在b 不包含时抛出错误任何一个词。

      你也可以看到这个问题:Get the first item from an iterable that matches a condition

      从评分最高的答案中应用该方法将给出

      new_df["c"]= new_df["b"].apply(lambda my_string: next(word for word in search_list if word.lower() in my_string.lower())
      

      【讨论】:

        【解决方案4】:

        您可以使用extract 并过滤掉那些nan(即不匹配):

        search_list = ['STEEL','IRON','GOLD','SILVER']
        
        df['c'] = df.b.str.extract('({0})'.format('|'.join(search_list)), flags=re.IGNORECASE)
        result = df[~pd.isna(df.c)]
        
        print(result)
        

        输出

                      a       b      c
        123 'Blah  Blah  Steel'  Steel
        789 'Blah  Blah   Gold'   Gold
        

        请注意,您必须导入 re 模块才能使用re.IGNORECASE 标志。作为替代方案,您可以直接使用2,这是re.IGNORECASE 标志的值。

        更新

        正如@user3483203 所述,您可以使用以下方法保存导入:

        df['c'] = df.b.str.extract('(?i)({0})'.format('|'.join(search_list)))
        

        【讨论】:

        • 你不需要额外的导入,df.b.str.extract('(?i)({0})'.format('|'.join(search_list))) 就可以了
        • 太棒了。谢谢
        • 如果第一条记录说,“Blah Blah Steel Gold”,您如何在 c 列中同时报告这两个词?我目前遇到了这个问题,我在搜索列表中点击了多个单词,并想记录所有这些单词,最好用逗号分隔。
        【解决方案5】:

        使用

        s=pd.DataFrame(df.b.str.upper().str.strip("'").str.split(' ').tolist())
        s.where(s.isin(search_list),'').sum(1)
        Out[492]: 
        0    STEEL
        1         
        2     GOLD
        dtype: object
        df['New']=s.where(s.isin(search_list),'').sum(1)
        df
        Out[494]: 
             a                  b    New
        0  123  'Blah Blah Steel'  STEEL
        1  456   'Blah Blah Blah'       
        2  789   'Blah Blah Gold'   GOLD
        

        【讨论】:

          【解决方案6】:

          这里,最终结果的解决方案就像你的显示:

          search_list = ['STEEL','IRON','GOLD','SILVER']
          
          def process(x):
              for s in search_list:
                  if s in x['b'].upper(): print("'"+ s +"'");return "'"+ s +"'"
              return ''
          
          df['c']= df.apply(lambda x: process(x),axis=1)
          df = df.drop(df[df['c'] == ''].index).reset_index(drop=True)
          
          print(df)
          

          输出:

               a                 b        c
          0  123  'Blah Blah Steel  'STEEL'
          1  789  'Blah Blah Gold'   'GOLD'
          

          【讨论】:

          • 有没有办法修改这个函数,使得 b 列的 search_list 中有多个单词,它们都将在 c 列中返回,比如逗号分隔?
          【解决方案7】:

          你可以使用:

          search_list = ['STEEL','IRON','GOLD','SILVER']
          pat = r'\b|\b'.join(search_list)
          pat2 = r'({})'.format('|'.join(search_list))
          
          df_new= df.loc[df.b.str.contains(pat,case=False,na=False)].reset_index(drop=True)
          df_new['new_col']=df_new.b.str.upper().str.extract(pat2)
          print(df_new)
          
               a                  b new_col
          0  123  'Blah Blah Steel'   STEEL
          1  789   'Blah Blah Gold'    GOLD
          

          【讨论】:

            猜你喜欢
            • 2020-03-16
            • 1970-01-01
            • 2017-07-27
            • 2020-03-12
            • 1970-01-01
            • 2020-02-26
            • 2016-08-12
            • 1970-01-01
            • 2021-02-04
            相关资源
            最近更新 更多