【问题标题】:Lemmatization of pandas column using Wordnet after POSPOS 后使用 Wordnet 对 pandas 列进行词形还原
【发布时间】:2020-01-03 10:25:02
【问题描述】:

我有一个带有文本的熊猫专栏 df_travail[line_text]。

我想对这一列的每个单词进行词形还原。

首先我将文本小写:

df_travail ['lowercase'] = df_travail['line_text'].str.lower()

然后,我对其进行标记并应用 POS(因为 wordnet 默认配置将每个单词都视为名词)。

from nltk import word_tokenize, pos_tag
tok_and_tag = lambda x: pos_tag(word_tokenize(x))
df_travail ['tok_and_tag'] = df_travail['lowercase'].apply(tok_and_tag)

然后我有以下内容:(整个df_travail['tok_and_tag']的摘录

"[('so', 'RB'), ('you', 'PRP'), (""'ve"", 'VBP'), ('come', 'VBN'), ('to', 'TO'), ('the', 'DT'), ('master', 'NN'), ('for', 'IN'), ('guidance', 'NN'), ('?', '.'), ('is', 'VBZ'), ('this', 'DT'), ('what', 'WP'), ('you', 'PRP'), (""'re"", 'VBP'), ('saying', 'VBG'), (',', ','), ('grasshopper', 'NN'), ('?', '.')]"
[('actually', 'RB'), (',', ','), ('you', 'PRP'), ('called', 'VBD'), ('me', 'PRP'), ('in', 'IN'), ('here', 'RB'), (',', ','), ('but', 'CC'), ('yeah', 'UH'), ('.', '.')]

但是,为了考虑到我应用了 POS 的事实,我对要应用的词形还原函数(使用 Wordnet)感到迷茫?

编辑:以下链接未提及我的问题的 POS 部分 Lemmatization of all pandas cells

【问题讨论】:

  • 不,我知道这篇文章,没有提到 POS...
  • 能否提供样本数据
  • 这能回答你的问题吗? wordnet lemmatization and pos tagging in python
  • 这应该有帮助:link
  • 好的,那么我知道我必须更改分类以使其不那么具体并与 wordnet lematizer 匹配。但是,如何将所有内容与我的 pandas 列混合?

标签: python pandas nltk wordnet lemmatization


【解决方案1】:

这是一个示例代码,用于考虑使用 'VERB' 而不是 'NOUN' 进行词形还原:

from nltk.stem import WordNetLemmatizer
wordnet_lemmatizer = WordNetLemmatizer()

def convert(text):
    lemmatized_text = []
    for i in text.split():
        lemmatized_text.append(str(wordnet_lemmatizer.lemmatize(i,pos="v")))

    return ' '.join(lemmatized_text)

df['text'] = df['text'].apply(lambda x: convert(x))

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2014-05-24
    • 2015-09-10
    • 1970-01-01
    • 2018-01-05
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多