【问题标题】:Spacy: count occurrence for specific token in each sentenceSpacy:计算每个句子中特定标记的出现次数
【发布时间】:2021-12-18 04:49:52
【问题描述】:
我想使用 spacy 计算语料库中每个句子的标记 和 的出现,并将每个句子的结果附加到列表中。到现在为止,下面的代码返回关于和的总数(整个语料库)。
3 个句子的示例/期望输出:['1', '0', '2']
当前输出:[3]
doc = nlp(corpus)
nb_and = []
for sent in doc.sents:
i = 0
for token in sent:
if token.text == "and":
i += 1
nb_and.append(i)
【问题讨论】:
标签:
nlp
counter
spacy
find-occurrences
spacy-3
【解决方案1】:
每句处理完后需要在nb_and后面追加i:
for sent in doc.sents:
i = 0
for token in sent:
if token.text == "and":
i += 1
nb_and.append(i)
测试代码:
import spacy
nlp = spacy.load("en_core_web_trf")
corpus = "I see a cat and a dog. None seems to be unhappy. My mother and I wanted to buy a parrot and a tortoise."
doc = nlp(corpus)
nb_and = []
for sent in doc.sents:
i = 0
for token in sent:
if token.text == "and":
i += 1
nb_and.append(i)
nb_and
# => [1, 0, 2]