【问题标题】:Finding pairs of rows with matching column sub-strings in pandas dataframe在熊猫数据框中查找具有匹配列子字符串的行对
【发布时间】:2019-10-17 22:13:59
【问题描述】:

我有一个包含几列的数据框。其中一个名为'log_text'. 我想在此列中查找具有匹配字符串的行对。

例如,如果'log_text' 有这些字符串

 Device remove ID#xxx  
 Device remove ID#yyy  
 Device remove ID#zzz  
 Device arrive ID#xxx  
 Device arrive ID#yyy 
 Device arrive ID#zzz 

目标: 我想获取包含'Device remove ID#xxx' 和'Device arrive ID#xxx' 的行并能够对它们的其他列进行处理,然后对包含'Device remove ID#yyy' 和'Device arrive ID#yyy' 等的行重复此操作。

我尝试的是使用iterrows(),找到当前行的ID#,从表中删除该行,然后找到包含匹配ID#字符串的第一行。

    for index, row in temp_df.iterrows():
        log_string = row['log_text']
        id_text = log_string.partition("ID#")[2]
        temp_df.drop(row)
        match = temp_df[temp_df['log_text'].str.contains(id_text)]
        # Somehow stash the 2 rows together somewhere? 
            # like stash[index,1] = row; stash[index,2] = match;
        temp_df.drop(match)

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    你可以使用pandas.Series.str.split和pandas.groupby:

    In [10]: df = pd.DataFrame({'log':['Device remove ID#xxx',
        ...:                           'Device remove ID#yyy',
        ...:                           'Device remove ID#zzz',
        ...:                           'Device arrive ID#xxx',
        ...:                           'Device arrive ID#yyy',
        ...:                           'Device arrive ID#zzz',],
                                'other_row':[1,2,3,42,54,6]})
    
    In [11]: df
    Out[11]:
                        log  other_row
    0  Device remove ID#xxx          1
    1  Device remove ID#yyy          2
    2  Device remove ID#zzz          3
    3  Device arrive ID#xxx         42
    4  Device arrive ID#yyy         54
    5  Device arrive ID#zzz          6
    
    In [14]: df_splits = df['log'].str.split(expand=True)
    
    In [16]: df['action'] = df_splits[1]
    
    In [17]: df['user'] = df_splits[2]
    
    In [18]: df
    Out[18]:
                        log  other_row  action    user
    0  Device remove ID#xxx          1  remove  ID#xxx
    1  Device remove ID#yyy          2  remove  ID#yyy
    2  Device remove ID#zzz          3  remove  ID#zzz
    3  Device arrive ID#xxx         42  arrive  ID#xxx
    4  Device arrive ID#yyy         54  arrive  ID#yyy
    5  Device arrive ID#zzz          6  arrive  ID#zzz
    
    
    In [22]: for i, d in df.groupby('user'):
        ...:     print i
        ...:     print d
        ...:     print d['other_row'].sum()
        ...:     print
        ...:
        ...:
    ID#xxx
                        log  other_row  action    user
    0  Device remove ID#xxx          1  remove  ID#xxx
    3  Device arrive ID#xxx         42  arrive  ID#xxx
    43
    
    ID#yyy
                        log  other_row  action    user
    1  Device remove ID#yyy          2  remove  ID#yyy
    4  Device arrive ID#yyy         54  arrive  ID#yyy
    56
    
    ID#zzz
                        log  other_row  action    user
    2  Device remove ID#zzz          3  remove  ID#zzz
    5  Device arrive ID#zzz          6  arrive  ID#zzz
    9
    

    【讨论】:

      【解决方案2】:

      IIUC,

      我觉得你可以使用.str.count和.loc做进一步的操作

      例如:

      rows_to_filter = ['Device remove ID#xxx','Device remove ID#yyy',
      'Device remove ID#zzz','Device arrive ID#xxx',
      'Device arrive ID#yyy','Device arrive ID#zzz']
      
      df.loc[df['log_text'].str.count('|'.join(rows_to_filter)) > 1, 'col'] = 'do something'
      

      这将返回一个数据帧切片,其中包含在任何给定行中出现多次以上列表的任何内容,您可能需要修改逻辑,因为如果没有示例输出,我不是 100% 需要的。

      【讨论】:

        【解决方案3】:

        如果您需要保留原始列并且只想对最后 3 个字符进行排序,则可以为此目的创建一个单独的列。

        df1['group'] = df1['log_text'].str[-3::]
        

        这将创建“log_text”列的副本,但只保留最后三个字符。

        【讨论】:

          猜你喜欢
          • 2021-11-15
          • 2020-05-31
          • 2018-04-26
          • 2021-10-07
          • 1970-01-01
          • 1970-01-01
          • 2017-11-17
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多