【发布时间】:2021-03-24 22:55:26
【问题描述】:
我想从一个网站上抓取的 cmets 列表中提取一组特定的单词来计算它们,并使用我的 TextBlob 词典中最常见的单词,这将用于简单的情绪分析。简化:我想得到所有可能有正面或负面情绪的形容词。 komentarze 是一个巨大的字符串列表,每个字符串都是一个句子,我想检查一下情绪。 我想从这个字符串列表中创建一个单词列表,然后检查哪些形容词最常见,这些形容词既不是标点符号也不是停用词,并且在动词之前。。当我运行我的代码时,我收到一个错误:IndexError: [E040] Attempt to access token at 18, max length 18. 这个错误代表 Attempt to access token at {i},最大长度 {max_length}。 我尝试了不同的代码,但都不起作用。
这是一个想要继续但给出 E040 错误的代码示例:
import spacy
import json
import pandas as pd
from spacy.lang.pl.stop_words import STOP_WORDS
from spacy.tokens import Token
from spacy.lang.pl.examples import sentences
from collections import Counter
with open('directory/file.json', mode='r') as f:
dane = json.load(f)
df = pd.DataFrame(dane)
komentarze = df['komentarz'].tolist()
nlp = spacy.load('pl_core_news_lg')
slowa = []
zwroty = []
for doc in nlp.pipe(komentarze):
#here I want to extract most common words
slowa += [token.text for token in doc if not token.is_stop and not token.is_punct]
#here I want to extract adjs, that are not puncts nor stop-words and are before a verb.
zwroty += [token.text for token in doc if (not token.is_stop and not token.is_punct and
token.pos_ == "ADJ" and doc[token.i + 1].pos_ == "VERB")]
zwroty_freq = Counter(zwroty)
common_zwroty = zwroty_freq.most_common(100)
print(common_zwroty)
当我在循环中运行一个额外的adjsy += [token.text for token in doc if (not token.is_stop and not token.is_punct and token.pos_ == "ADJ")] 时,一切正常,但我根本无法指定ADJ 之前或之后的单词。
我可以通过以下方式遍历一个简单的字符串:
for token in doc:
if token.pos_ == 'ADJ':
if doc[token.i + 1].pos_ == 'VERB':
print('yaaay’)
但我真的不知道如何在我的循环中设置它。 我也试过了:
for token in doc:
if not token.is_stop and not token.is_punct:
if token.pos_ == "ADJ":
if doc[token.i+1].pos_ == "NOUN" in range(1):
zwroty += token.text
但这只给了我字母。
如何解决我的问题以获得我想要的结果?
这个循环中的文本是否也可能越低?我试了好几次,都没有成功……
已编辑:
我按照@polm23 的提议修改了我的代码。好吧,它有效,但我无法将这个方法与我的 [w.lemma_ for w in doc if not w.is_stop and not w.is_punct and not w.like_num and w.pos_ == "VERB"] 列表理解一起加入,这给了我一个错误:ValueError: [E195] Matcher can be called on Doc or Span only, got Token.
感谢@polm23,这是一段代码,它可以工作,但考虑到了帐号、标点符号等:
import everything I need
with open('file.json', mode='r') as f:
dane = json.load(f)
df = pd.DataFrame(dane)
komentarze = df['komentarz'].tolist()
nlp = spacy.load('pl_core_news_lg')
matcher = Matcher(nlp.vocab)
patterns = [[{'POS':'ADJ'}, {'POS':'NOUN'}]]
matcher.add("demo", patterns)
zwroty =[]
for doc in nlp.pipe(komentarze):
matches = matcher(doc)
for match_id, start, end in matches:
string_id = nlp.vocab.strings[match_id]
span = doc[start:end]
zwroty += (match_id, string_id, start, end, span.text)
这是一段代码,它不起作用,但是,应该考虑到这一点:
for w in doc:
if not w.is_stop and not w.is_punct:
w.lemma_
matches = matcher(w)
for match_id, start, end in matches:
string_id = nlp.vocab.strings[match_id]
span = w[start:end]
zwroty += (match_id, string_id, start, end, span.text)
【问题讨论】:
标签: python extract spacy tokenize