【问题标题】:Singular and Plural words matching with Pandas与 Pandas 匹配的单复数单词
【发布时间】:2015-09-13 06:17:01
【问题描述】:

这个问题是我之前的问题Multiple Phrases Matching Python Pandas 的延伸。虽然我在得到解决问题的答案后想出了办法,但还是出现了一些典型的单复数问题。

ingredients=pd.Series(["vanilla extract","walnut","oat","egg","almond","strawberry"])

df=pd.DataFrame(["1 teaspoons vanilla extract","2 eggs","3 cups chopped walnuts","4 cups rolled oats","1 (10.75 ounce) can Campbell's Condensed Cream of Chicken with Herbs Soup","6 ounces smoke-flavored almonds, finely chopped","sdfgsfgsf","fsfgsgsfgfg","2 small strawberries"])

我只需要将成分系列中的短语与 DataFrame 中的短语相匹配。作为伪代码,

如果在 DataFrame 中的短语中找到成分(单数或复数), 退回成分。否则,返回 false。

这是通过以下给出的答案实现的,

df.columns = ['val']
V = df.val.str.lower().values.astype(str)
K = ingredients.values.astype(str)
df['existence'] = map(''.join, np.where(np.char.count(V, K[...,np.newaxis]),K[...,np.newaxis], '').T)

我还应用了以下方法来用 NAN 填充空单元格,以便我可以轻松过滤掉数据。

df.ix[df.existence=='', 'existence'] = np.nan

我们的结果如下,

print df
                                                 val        existence
0                        1 teaspoons vanilla extract  vanilla extract
1                                             2 eggs              egg
2                             3 cups chopped walnuts           walnut
3                                 4 cups rolled oats              oat
4  1 (10.75 ounce) can Campbell's Condensed Cream...             NaN    
5    6 ounces smoke-flavored almonds, finely chopped           almond
6                                          sdfgsfgsf              NaN  
7                                        fsfgsgsfgfg              NaN
8  2 small strawberries                                           NaN

这一直是正确的,但是当单数和复数单词映射不像almond=> almonds apple=> apples。当出现strawberry=>strawberries 之类的内容时,此代码会将其识别为NaN。

改进我的代码以检测此类事件。我喜欢将我的成分 Series 更改为 data Frame 如下。

#ingredients

#inputwords       #outputword

vanilla extract    vanilla extract 
walnut             walnut
walnuts            walnut
oat                oat
oats               oat
egg                egg
eggs               egg
almond             almond
almonds            almond
strawberry         strawberry
strawberries       strawberry
cherry             cherry
cherries           cherry

所以我的逻辑是每当#inputwords 中的单词出现在短语中时,我想返回另一个单元格中的单词。换句话说,当strawberry 或strawberries 出现在短语中时,刚刚输出的代码将在它旁边放上strawberry。这样我的最终结果将是

                                                 val        existence
0                        1 teaspoons vanilla extract  vanilla extract
1                                             2 eggs              egg
2                             3 cups chopped walnuts           walnut
3                                 4 cups rolled oats              oat
4  1 (10.75 ounce) can Campbell's Condensed Cream...             NaN    
5    6 ounces smoke-flavored almonds, finely chopped           almond
6                                          sdfgsfgsf              NaN  
7                                        fsfgsgsfgfg              NaN
8  2 small strawberries                                           strawberry

我找不到将此功能合并到现有代码或编写新代码的方法。谁能帮我解决这个问题?

【问题讨论】:

    标签: python regex pandas


    【解决方案1】:

    考虑使用词干分析器 :) http://www.nltk.org/howto/stem.html

    直接从他们的页面中取出:

        from nltk.stem.snowball import SnowballStemmer
        stemmer = SnowballStemmer("english")
        stemmer2 = SnowballStemmer("english", ignore_stopwords=True)
        >>> print(stemmer.stem("having"))
        have
        >>> print(stemmer2.stem("having"))
        having
    

    在将句子中的所有单词与成分列表匹配之前,重构您的代码。

    nltk 是一款非常棒的工具,可以满足您的要求!

    干杯

    【讨论】:

      【解决方案2】:
      # your data frame
      df = pd.DataFrame(data = ["1 teaspoons vanilla extract","2 eggs","3 cups chopped walnuts","4 cups rolled oats","1 (10.75 ounce) can Campbell's Condensed Cream of Chicken with Herbs Soup","6 ounces smoke-flavored almonds, finely chopped","sdfgsfgsf","fsfgsgsfgfg","2 small strawberries"])
      
      # Here you create mapping
      mapping = pd.Series(index = ['vanilla extract' , 'walnut','walnuts','oat','oats','egg','eggs','almond','almonds','strawberry','strawberries','cherry','cherries'] , 
                data = ['vanilla extract' , 'walnut','walnut','oat','oat','egg','egg','almond','almond','strawberry','strawberry','cherry','cherry'])
      # create a function that checks if the value you're looking for exist in specific phrase or not
      def get_match(df):
          match = np.nan
          for key , value in mapping.iterkv():
              if key in df[0]:
                  match = value
          return match
      # apply this function on each row
      df.apply(get_match, axis = 1)
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2015-11-23
        • 2016-02-17
        • 1970-01-01
        • 1970-01-01
        • 2012-03-23
        • 1970-01-01
        相关资源
        最近更新 更多