【问题标题】:How to call stopwords from a txt file using wordcloud STOPWORDS如何使用 wordcloud STOPWORDS 从 txt 文件中调用停用词
【发布时间】:2019-08-18 08:20:48
【问题描述】:

我正在从 pdf 文件中提取 wordcloud。我可以从列表中提取停用词,但不能使用 txt 文件提取。我知道调用文件的路径有问题。

我已成功使用列表编辑停用词,但我希望能够将 txt 文件用于停用词,因为最终我想将不同的停用词文件关联用于不同目的。

提前感谢您的帮助。

#viz libs
from wordcloud import WordCloud, STOPWORDS
import matplotlib.pyplot as plt
#img libs
from PIL import Image
#binary array lib
import numpy as np
#pdf reader
import PyPDF4

pdfFileObj = open('Test-Resume-Doc.pdf', 'rb')
pdfReader = PyPDF4.PdfFileReader(pdfFileObj)
print(pdfReader.numPages)
pageObj = pdfReader.getPage(0)
pageText = (pageObj.extractText())
pdfFileObj.close()
#set stopwords
stopwords = set(STOPWORDS)

#can call stopwords from a list as such
#stopwords.update(["word1", "word2", "word3", ...])
#call stopwords from txt file and program executes ignoring txt file, the problem is how the path is run
stopwords.update(['stopwords.txt'])

rsMask = np.array(Image.open('Resume_WordCloud.png'))
#create wordcloud with stopwords
cloud = WordCloud(stopwords=stopwords, background_color="black", mask=rsMask).generate(pageObj.extractText())


plt.imshow(cloud, interpolation="bilinear")
plt.axis("off")
plt.savefig('path.../PythonPDFRW/Resume_WordCloud_fromPython.png'.format(cloud))
plt.show()```

【问题讨论】:

  • 好的,我已经尝试将文本读入 readline,所以下一步是拆分文本?目前它将文本文件作为单个长字符串读取...text_file = open("stopwords.txt", "r")stopwords = set(STOPWORDS)stopwords.update(text_file.readlines())

标签: python-3.x stop-words word-cloud


【解决方案1】:

使用 for 循环读取文件中的每一行,然后去掉换行符 \n 它并不优雅,但很有效。

for line in text_file:
    stopwords.add(line.strip('\n'))
print(stopwords)```

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2019-05-28
    • 1970-01-01
    • 2020-09-09
    • 2019-10-01
    • 2017-05-21
    • 2013-06-23
    • 1970-01-01
    相关资源
    最近更新 更多