【问题标题】:how to resolve the error: AttributeError: 'generator' object has no attribute 'endswith'如何解决错误:AttributeError: 'generator' object has no attribute 'endswith'
【发布时间】:2018-01-19 12:19:17
【问题描述】:

当我尝试运行此代码来预处理文本时,我收到以下错误,有人遇到了类似的问题,但帖子没有足够的详细信息。

我在这里将所有内容都放在上下文中,希望能帮助审阅者更好地帮助我们。

这里是函数;

def preprocessing(text):
    #text=text.decode("utf8")
    #tokenize into words
    tokens=[word for sent in nltk.sent_tokenize(text) for word in 
    nltk.word_tokenize(sent)]
    #remove stopwords
    stop=stopwords.words('english')
    tokens=[token for token in tokens if token not in stop]
    #remove words less than three letters
    tokens=[word for word in tokens if len(word)>=3]
    #lower capitalization
    tokens=[word.lower() for word in tokens]
    #lemmatization
    lmtzr=WordNetLemmatizer()
    tokens=[lmtzr.lemmatize(word for word in tokens)]
    preprocessed_text=' '.join(tokens)
    return preprocessed_text

在这里调用函数;

#open the text data from disk location
sms=open('C:/Users/Ray/Documents/BSU/Machine_learning/Natural_language_Processing_Pyhton_And_NLTK_Chap6/smsspamcollection/SMSSpamCollection')
sms_data=[]
sms_labels=[]
csv_reader=csv.reader(sms,delimiter='\t')
for line in csv_reader:
    #adding the sms_id
    sms_labels.append(line[0])
    #adding the cleaned text by calling the preprocessing method
    sms_data.append(preprocessing(line[1]))
sms.close()

结果;

--------------------------------------------------------------------------- AttributeError                            Traceback (most recent call last) <ipython-input-38-b42d443adaa6> in <module>()
      8     sms_labels.append(line[0])
      9     #adding the cleaned text by calling the preprocessing method
---> 10     sms_data.append(preprocessing(line[1]))
     11 sms.close()

<ipython-input-37-69ef4cd83745> in preprocessing(text)
     12     #lemmatization
     13     lmtzr=WordNetLemmatizer()
---> 14     tokens=[lmtzr.lemmatize(word for word in tokens)]
     15     preprocessed_text=' '.join(tokens)
     16     return preprocessed_text

~\Anaconda3\lib\site-packages\nltk\stem\wordnet.py in lemmatize(self, word, pos)
     38 
     39     def lemmatize(self, word, pos=NOUN):
---> 40         lemmas = wordnet._morphy(word, pos)
     41         return min(lemmas, key=len) if lemmas else word
     42 

~\Anaconda3\lib\site-packages\nltk\corpus\reader\wordnet.py in
_morphy(self, form, pos, check_exceptions)    1798     1799         # 1. Apply rules once to the input to get y1, y2, y3, etc.
-> 1800         forms = apply_rules([form])    1801     1802         # 2. Return all that are in the database (and check the original too)

~\Anaconda3\lib\site-packages\nltk\corpus\reader\wordnet.py in apply_rules(forms)    1777         def apply_rules(forms):    1778     return [form[:-len(old)] + new
-> 1779                     for form in forms    1780                     for old, new in substitutions    1781                     if form.endswith(old)]

~\Anaconda3\lib\site-packages\nltk\corpus\reader\wordnet.py in <listcomp>(.0)    1779                     for form in forms    1780   for old, new in substitutions
-> 1781                     if form.endswith(old)]    1782     1783         def filter_forms(forms):

AttributeError: 'generator' object has no attribute 'endswith'

我相信错误来自 nltk.corpus.reader.wordnet 的源代码

整个源代码可以在 nltk 文档页面中看到。在这里发帖太长了;但下面是原始的link:

感谢您的帮助。

【问题讨论】:

  • 你在这里传递了一个生成器:tokens=[lmtzr.lemmatize(word for word in tokens)] - 这个方法真的接受一个生成器吗?你是说tokens=[lmtzr.lemmatize(word) for word in tokens]
  • 嘿,请看这个stackoverflow.com/questions/47769818/… 并理解为什么你的预处理不是最佳的,因为它会多次迭代令牌......请问你从哪里得到这个代码示例?此代码示例令人困扰,您不是第一个基于此提出问题的人。
  • 是的,代码来自 Hardeniya 等人的《Natural Language Processing Python And NLTK》一书的第 6 章(文本分类)。
  • 是的,代码来自 Hardeniya 等人的《Natural Language Processing Python And NLTK》一书的第 6 章(文本分类)。

标签: python nltk preprocessor wordnet lemmatization


【解决方案1】:

错误消息和回溯将您指向问题的根源:

在预处理中(文本) 12 #lemmatization 13 lmtzr=WordNetLemmatizer() ---> 14 tokens=[lmtzr.lemmatize(word for word in tokens)] 15 preprocessed_text=' '.join(tokens) 16 return preprocessed_text

~\Anaconda3\lib\site-packages\nltk\stem\wordnet.py in lemmatize(self, word, pos) 38 39 def lemmatize(self, word, pos=NOUN):

显然,从函数的签名(word,而不是words)和错误(“没有属性'endswith'”-endswith() 实际上是str 方法),lemmatize() 期望单个字,但在这里:

tokens=[lmtzr.lemmatize(word for word in tokens)]

你正在传递一个生成器表达式。

你想要的是:

tokens = [lmtzr.lemmatize(word) for word in tokens]

注意:您提到:

我相信错误来自源代码 nltk.corpus.reader.wordnet

错误确实是在这个包中引发的,但它“来自”(在“由”的意义上)您的代码传递了错误的参数;)

希望这对您下次自己调试此类问题有所帮助。

【讨论】:

  • str.endswith() 接受一个元组,这会像lmtzr.lemmatize(tuple(tokens))] 一样在这里工作吗?
  • @Chris_Rands 显然不是。重新阅读调用堆栈并查看str.endswith() 是如何被调用的(提示:这是一个在str 实例上调用的方法——实际上是传递给.lemmatize() 的单词)。如果您传递了tuple,则将在该元组上调用endswith()(就像在OP 代码中的生成器上调用它一样),并引发相同的AttributeError,因为tuple 也没有endswith() 属性。 TL;DR : .lemmatize() 需要一个字符串,给它一个字符串,否则它会失败。
  • 啊,是的,点了,我应该更仔细地看,同意!
  • 感谢 Bruno 的输入,问题出在参数上。
猜你喜欢
  • 2021-08-22
  • 2021-05-16
  • 1970-01-01
  • 1970-01-01
  • 2021-04-06
  • 2022-12-04
  • 2022-11-11
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多