您将文本设置为 spaCy 类型是正确的 - 您希望将每个标记元组转换为 spaCy Doc。从那里,最好使用标记的属性来回答“标记是停用词”(使用token.is_stop)或“这个标记的引理是什么”(使用token.lemma_)的问题。我的实现如下,我稍微更改了您的输入数据以包含一些复数示例,以便您可以看到词形还原正常工作。
import spacy
import pandas as pd
nlp = spacy.load('en_core_web_sm')
texts = [('the','cheeseburger','was','great'),
('i','never','did','like','the','pizzas','too','much'),
('yellowed','submarines','was','only','an','ok','song')]
df = pd.DataFrame({'word_tokens': texts})
初始DataFrame如下所示:
|
word_tokens |
| 0 |
('the', 'cheeseburger', 'was', 'great') |
| 1 |
('i', 'never', 'did', 'like', 'the', 'pizzas', 'too', 'much') |
| 2 |
('yellowed', 'submarines', 'was', 'only', 'an', 'ok', 'song') |
我定义了执行主要任务的函数:
- token 元组 -> spaCy Doc
- spaCy Doc -> 非停用词列表
- spaCy Doc -> 不间断的词形还原词列表
def to_doc(words:tuple) -> spacy.tokens.Doc:
# Create SpaCy documents by joining the words into a string
return nlp(' '.join(words))
def remove_stops(doc) -> list:
# Filter out stop words by using the `token.is_stop` attribute
return [token.text for token in doc if not token.is_stop]
def lemmatize(doc) -> list:
# Take the `token.lemma_` of each non-stop word
return [token.lemma_ for token in doc if not token.is_stop]
应用这些看起来像:
# create documents for all tuples of tokens
docs = list(map(to_doc, df.word_tokens))
# apply removing stop words to all
df['removed_stops'] = list(map(remove_stops, docs))
# apply lemmatization to all
df['lemmatized'] = list(map(lemmatize, docs))
你得到的输出应该是这样的:
|
word_tokens |
removed_stops |
lemmatized |
| 0 |
('the', 'cheeseburger', 'was', 'great') |
['cheeseburger', 'great'] |
['cheeseburger', 'great'] |
| 1 |
('i', 'never', 'did', 'like', 'the', 'pizzas', 'too', 'much') |
['like', 'pizzas'] |
['like', 'pizza'] |
| 2 |
('yellowed', 'submarines', 'was', 'only', 'an', 'ok', 'song') |
['yellowed', 'submarines', 'ok', 'song'] |
['yellow', 'submarine', 'ok', 'song'] |
根据您的用例,您可能想要探索 spaCy 文档对象 (https://spacy.io/api/doc) 的其他属性。特别是,如果您想从文本中提取更多含义,请查看 doc.noun_chunks 和 doc.ents。
还值得注意的是,如果您打算将其用于大量文本,则应考虑nlp.pipe:https://spacy.io/usage/processing-pipelines。它分批而不是一个一个地处理您的文档,并且可以提高您的实施效率。