【问题标题】:SpaCy extraction of an adjective, that precede a verb and isn't a stop word nor a punctuationSpaCy 提取形容词,在动词之前,既不是停用词也不是标点符号
【发布时间】:2021-03-24 22:55:26
【问题描述】:

我想从一个网站上抓取的 cmets 列表中提取一组特定的单词来计算它们,并使用我的 TextBlob 词典中最常见的单词,这将用于简单的情绪分析。简化:我想得到所有可能有正面或负面情绪的形容词。 komentarze 是一个巨大的字符串列表,每个字符串都是一个句子,我想检查一下情绪。 我想从这个字符串列表中创建一个单词列表,然后检查哪些形容词最常见,这些形容词既不是标点符号也不是停用词,并且在动词之前。。当我运行我的代码时,我收到一个错误:IndexError: [E040] Attempt to access token at 18, max length 18. 这个错误代表 Attempt to access token at {i},最大长度 {max_length}。 我尝试了不同的代码,但都不起作用。

这是一个想要继续但给出 E040 错误的代码示例:

import spacy
import json
import pandas as pd
from spacy.lang.pl.stop_words import STOP_WORDS
from spacy.tokens import Token
from spacy.lang.pl.examples import sentences
from collections import Counter

with open('directory/file.json', mode='r') as f:
    dane = json.load(f)

df = pd.DataFrame(dane)
komentarze = df['komentarz'].tolist()

nlp = spacy.load('pl_core_news_lg')
slowa = []
zwroty = []

for doc in nlp.pipe(komentarze):
    #here I want to extract most common words
    slowa += [token.text for token in doc if not token.is_stop and not token.is_punct]
    #here I want to extract adjs, that are not puncts nor stop-words and are before a verb.
    zwroty += [token.text for token in doc if (not token.is_stop and not token.is_punct and 
    token.pos_ == "ADJ" and doc[token.i + 1].pos_ == "VERB")]

zwroty_freq = Counter(zwroty)
common_zwroty = zwroty_freq.most_common(100)
print(common_zwroty)

当我在循环中运行一个额外的adjsy += [token.text for token in doc if (not token.is_stop and not token.is_punct and token.pos_ == "ADJ")] 时,一切正常,但我根本无法指定ADJ 之前或之后的单词。

我可以通过以下方式遍历一个简单的字符串:

for token in doc:
    if token.pos_ == 'ADJ':
        if doc[token.i + 1].pos_ == 'VERB':
            print('yaaay’)

但我真的不知道如何在我的循环中设置它。 我也试过了:

   for token in doc:
        if not token.is_stop and not token.is_punct:
            if token.pos_ == "ADJ":
                if doc[token.i+1].pos_ == "NOUN" in range(1):
                    zwroty += token.text

但这只给了我字母。

如何解决我的问题以获得我想要的结果?

这个循环中的文本是否也可能越低?我试了好几次,都没有成功……

已编辑: 我按照@polm23 的提议修改了我的代码。好吧,它有效,但我无法将这个方法与我的 [w.lemma_ for w in doc if not w.is_stop and not w.is_punct and not w.like_num and w.pos_ == "VERB"] 列表理解一起加入,这给了我一个错误:ValueError: [E195] Matcher can be called on Doc or Span only, got Token.

感谢@polm23,这是一段代码,它可以工作,但考虑到了帐号、标点符号等:

import everything I need

with open('file.json', mode='r') as f:
    dane = json.load(f)

df = pd.DataFrame(dane)
komentarze = df['komentarz'].tolist()

nlp = spacy.load('pl_core_news_lg')
matcher = Matcher(nlp.vocab)
patterns = [[{'POS':'ADJ'}, {'POS':'NOUN'}]]
matcher.add("demo", patterns)

zwroty =[]

for doc in nlp.pipe(komentarze):
    matches = matcher(doc)

    for match_id, start, end in matches:
        string_id = nlp.vocab.strings[match_id]
        span = doc[start:end]
        zwroty += (match_id, string_id, start, end, span.text)

这是一段代码,它不起作用,但是,应该考虑到这一点:

    for w in doc:
        if not w.is_stop and not w.is_punct:
            w.lemma_

            matches = matcher(w)
            for match_id, start, end in matches:
                string_id = nlp.vocab.strings[match_id]
                span = w[start:end]
                zwroty += (match_id, string_id, start, end, span.text)

【问题讨论】:

    标签: python extract spacy tokenize


    【解决方案1】:

    这是 spaCy 的 Matchers 的完美用例。下面是一个匹配英文 ADJ NOUN 的例子:

    import spacy
    from spacy.matcher import Matcher
    
    nlp = spacy.load("en_core_web_sm")
    
    matcher = Matcher(nlp.vocab)
    
    patterns = [
        [{'POS':'ADJ'}, {'POS':'NOUN'}],
        ]
    matcher.add("demo", patterns)
    
    doc = nlp("There is a red card in the blue envelope.")
    matches = matcher(doc)
    for match_id, start, end in matches:
        string_id = nlp.vocab.strings[match_id]  # Get string representation
        span = doc[start:end]  # The matched span
        print(match_id, string_id, start, end, span.text)
    

    输出:

    2193290520773312886 demo 3 5 red card
    2193290520773312886 demo 7 9 blue envelope
    

    您可以在Counter 中使用这些匹配项,或者根据需要跟踪频率。您还可以设置一个函数,以便在匹配时运行。

    这个循环中的文本是否也可能越低?试了好几次都没有效果……

    不完全确定你想做什么,但如果你有一个匹配功能,你可以在一个令牌上使用lower_ 属性。还可以看看lemma_,这可能会更好,尤其是动词。


    我不完全确定我理解你想要做什么,但看起来问题是你试图过滤令牌,然后将它们传递给 Matcher。相反,请在文档上使用 Matcher,然后过滤其输出。

    另外,标点符号永远不能是形容词,我不知道你为什么要检查它。

    out = []
    for doc in docs:
        matches = matcher(doc)
        # because we are just matching [ADJ NOUN] we know the first token is ADJ
        for match_id, start, end in matches:
            string_id = nlp.vocab.strings[match_id]
            adj = doc[start]
            # ignore stop words
            if adj.is_stop: continue
            # get the lemma
            lemma = adj.lemma_
    
            out += (adj, lemma) # or whatever
    

    【讨论】:

    • 谢谢你的回复,我没想到matcher!我已经尝试过你的模式,但不知何故我无法将它循环到我的 for doc in nlp.pipe(komentarze) 上然后创建一个 zwroty += (match_id, string_id, start, end, span.text) 对象,我将在该对象上使用计数器,我只得到一个错误,即 'list' object is not callable。
    • 呃,也许将您的代码添加到您的问题中?我不清楚你做错了什么。
    • 我编辑了我的问题 - 您的解决方案有效,但我不知道如何将它与词形还原以及清除文本所需的所有功能结合起来。
    • 这很好,我看到了我的错误,非常感谢! :)
    猜你喜欢
    • 2012-11-11
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-01-07
    • 1970-01-01
    • 2011-07-29
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多