【问题标题】:Pandas - check if a string column in one dataframe contains a pair of strings from another dataframePandas - 检查一个数据帧中的字符串列是否包含来自另一个数据帧的一对字符串
【发布时间】:2017-09-12 14:05:36
【问题描述】:

这个问题是基于我提出的另一个问题,我没有完全涵盖这个问题:Pandas - check if a string column contains a pair of strings

这是问题的修改版本。

我有两个数据框:

df1 = pd.DataFrame({'consumption':['squirrel ate apple', 'monkey likes apple', 
                                  'monkey banana gets', 'badger gets banana', 'giraffe eats grass', 'badger apple loves', 'elephant is huge', 'elephant eats banana tree', 'squirrel digs in grass']})

df2 = pd.DataFrame({'food':['apple', 'apple', 'banana', 'banana'], 
                   'creature':['squirrel', 'badger', 'monkey', 'elephant']})

目标是测试 df1.consumptions 中是否存在 df.food:df.creature 对。

上例中该测试的预期答案是:

['True', 'False', 'True', 'False', 'False', 'True', 'False', 'True', 'False']

模式是:

松鼠吃了苹果=真,因为松鼠和苹果是一对。 猴子喜欢苹果 = False,因为猴子和苹果不是我们要找的一对。

我正在考虑构建一个包含对值数据帧的字典,其中每个数据帧将用于一个生物,例如松鼠、猴子等,然后使用 np.where 创建一个布尔表达式并执行 str.contains。

不确定这是否是最简单的方法。

【问题讨论】:

    标签: python pandas string-matching boolean-expression


    【解决方案1】:

    这是我使用理解和zip的答案
    注意,这会检查df1中的子字符串

    c = df1.consumption.values.tolist()
    f = df2.food.values.tolist()
    a = df2.creature.values.tolist() 
    
    check = np.array([[fd in cs and cr in cs for fd, cr in zip(f, a)] for cs in c])
    
    check.any(1)
    
    array([ True, False,  True, False, False,  True, False,  True, False], dtype=bool)
    

    这是@MaxU 所做的pandas 版本。尊重他的所作所为……太棒了!

    X = df1.consumption.str.get_dummies(' ')
    Y = (df2.creature + ' ' + df2.food).str.get_dummies(' ') \
        .reindex_axis(X.columns, 1, fill_value=0)
    
    # This is where you can see which rows from `df2` (columns)
    # matched with which rows from `df1` (rows) 
    XY = X.dot(Y.T)
    
    print(XY)
    
       0  1  2  3
    0  2  1  0  0
    1  1  1  1  0
    2  0  0  2  1
    3  0  1  1  1
    4  0  0  0  0
    5  1  2  0  0
    6  0  0  0  1
    7  0  0  1  2
    8  1  0  0  0
    
    # return the desired `True`s and `False`s
    
    XY.gt(1).any(1)
    
    0     True
    1    False
    2     True
    3    False
    4    False
    5     True
    6    False
    7     True
    8    False
    dtype: bool
    

    幼稚测试

    【讨论】:

      【解决方案2】:

      考虑这种矢量化方法:

      from sklearn.feature_extraction.text import CountVectorizer
      
      vect = CountVectorizer()
      
      X = vect.fit_transform(df1.consumption)
      Y = vect.transform(df2.creature + ' ' + df2.food)
      
      res = np.ravel(np.any((X.dot(Y.T) > 1).todense(), axis=1))
      

      结果:

      In [67]: res
      Out[67]: array([ True, False,  True, False, False,  True, False,  True, False], dtype=bool)
      

      解释:

      In [68]: pd.DataFrame(X.toarray(), columns=vect.get_feature_names())
      Out[68]:
         apple  ate  badger  banana  digs  eats  elephant  gets  giraffe  grass  huge  in  is  likes  loves  monkey  squirrel  tree
      0      1    1       0       0     0     0         0     0        0      0     0   0   0      0      0       0         1     0
      1      1    0       0       0     0     0         0     0        0      0     0   0   0      1      0       1         0     0
      2      0    0       0       1     0     0         0     1        0      0     0   0   0      0      0       1         0     0
      3      0    0       1       1     0     0         0     1        0      0     0   0   0      0      0       0         0     0
      4      0    0       0       0     0     1         0     0        1      1     0   0   0      0      0       0         0     0
      5      1    0       1       0     0     0         0     0        0      0     0   0   0      0      1       0         0     0
      6      0    0       0       0     0     0         1     0        0      0     1   0   1      0      0       0         0     0
      7      0    0       0       1     0     1         1     0        0      0     0   0   0      0      0       0         0     1
      8      0    0       0       0     1     0         0     0        0      1     0   1   0      0      0       0         1     0
      
      In [69]: pd.DataFrame(Y.toarray(), columns=vect.get_feature_names())
      Out[69]:
         apple  ate  badger  banana  digs  eats  elephant  gets  giraffe  grass  huge  in  is  likes  loves  monkey  squirrel  tree
      0      1    0       0       0     0     0         0     0        0      0     0   0   0      0      0       0         1     0
      1      1    0       1       0     0     0         0     0        0      0     0   0   0      0      0       0         0     0
      2      0    0       0       1     0     0         0     0        0      0     0   0   0      0      0       1         0     0
      3      0    0       0       1     0     0         1     0        0      0     0   0   0      0      0       0         0     0
      

      更新:

      In [92]: df1['match'] = np.ravel(np.any((X.dot(Y.T) > 1).todense(), axis=1))
      
      In [93]: df1
      Out[93]:
                       consumption  match
      0         squirrel ate apple   True
      1         monkey likes apple  False
      2         monkey banana gets   True
      3         badger gets banana  False
      4         giraffe eats grass  False
      5         badger apple loves   True
      6           elephant is huge  False
      7  elephant eats banana tree   True
      8     squirrel digs in grass  False
      9        squirrel.eats/apple   True   # <----- NOTE
      

      【讨论】:

      • 谢谢 - 有一个警告 - 生物和食物的出现没有固定的模式。所以这个:Y = vect.transform(df2.creature + ' ' + df2.food) 不起作用。抱歉,我刚刚修改了问题中的消耗值以反映这一点。
      • @vagabond,您是否针对修改后的数据集测试了我的解决方案? ;-)
      • 哇!有用!有没有办法从它匹配的行中提取生物和食物?我的查找数据也在 100K ++ 行中。稀疏矩阵可能会拍内存?
      • @vagabond,稀疏矩阵非常节省内存(这是它们的主要目的)。关于提取 - 你能打开一个新问题并举个例子吗?
      • 嗯,是的,我可以-另一个问题确实是有时文本没有空格-squirrel.eats/apple。 . .它是 URL 数据。
      猜你喜欢
      • 2022-10-14
      • 2020-07-19
      • 1970-01-01
      • 2013-03-13
      • 2021-10-27
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多