【问题标题】:Python text processing: NLTK and pandasPython 文本处理:NLTK 和 pandas
【发布时间】:2016-04-19 10:47:08
【问题描述】:

我正在寻找一种在 Python 中构建可与额外数据一起使用的术语文档矩阵的有效方法。

我有一些带有其他属性的文本数据。我想对文本进行一些分析,并且希望能够将从文本中提取的特征(例如单个单词标记或 LDA 主题)与其他属性相关联。

我的计划是将数据加载为 pandas 数据框,然后每个响应将代表一个文档。不幸的是,我遇到了一个问题:

import pandas as pd
import nltk

pd.options.display.max_colwidth = 10000

txt_data = pd.read_csv("data_file.csv",sep="|")
txt = str(txt_data.comment)
len(txt)
Out[7]: 71581 

txt = nltk.word_tokenize(txt)
txt = nltk.Text(txt)
txt.count("the")
Out[10]: 45

txt_lines = []
f = open("txt_lines_only.txt")
for line in f:
    txt_lines.append(line)

txt = str(txt_lines)
len(txt)
Out[14]: 1668813

txt = nltk.word_tokenize(txt)
txt = nltk.Text(txt)
txt.count("the")
Out[17]: 10086

请注意,在这两种情况下,文本的处理方式都是只处理除空格、字母和 ,.?! 之外的任何内容。已删除(为简单起见)。

如您所见,转换为字符串的 pandas 字段返回的匹配项更少,字符串的长度也更短。

上面的代码有什么办法改进吗?

另外,str(x) 从 cmets 中创建了 1 个大字符串,而 [str(x) for x in txt_data.comment] 创建了一个无法分解为单词袋的列表对象。生成将保留文档索引的nltk.Text 对象的最佳方法是什么?换句话说,我正在寻找一种方法来创建术语文档矩阵,R 相当于来自 tm 包的 TermDocumentMatrix()。

非常感谢。

【问题讨论】:

  • 不确定您的问题是什么,但还有其他 NLP 库可能对您有所帮助,例如 pattern、textblob、C&C 等库,如果您遇到了死胡同,您也可以尝试这些库,他们每个人都比其他人有自己的优势。
  • 感谢@mid,我知道 gensim,但我以前从未听说过 textblob,但它确实看起来很有用!我对 Python 很陌生(我通常在 R 中工作),我真的怀疑我在 NLTK 上已经走到了死胡同,考虑到这个包有多受欢迎,我确定我只是错过了一些东西。跨度>

标签: python pandas machine-learning nltk


【解决方案1】:

使用pandasDataFrame 的好处是将nltk 功能应用于每个row,如下所示:

word_file = "/usr/share/dict/words"
words = open(word_file).read().splitlines()[10:50]
random_word_list = [[' '.join(np.random.choice(words, size=1000, replace=True))] for i in range(50)]

df = pd.DataFrame(random_word_list, columns=['text'])
df.head()

                                                text
0  Aaru Aaronic abandonable abandonedly abaction ...
1  abampere abampere abacus aback abalone abactor...
2  abaisance abalienate abandonedly abaff abacina...
3  Ababdeh abalone abac abaiser abandonable abact...
4  abandonable abandon aba abaiser abaft Abama ab...

len(df)

50

txt = df.text.apply(word_tokenize)
txt.head()

0    [Aaru, Aaronic, abandonable, abandonedly, abac...
1    [abampere, abampere, abacus, aback, abalone, a...
2    [abaisance, abalienate, abandonedly, abaff, ab...
3    [Ababdeh, abalone, abac, abaiser, abandonable,...
4    [abandonable, abandon, aba, abaiser, abaft, Ab...

txt.apply(len)

0     1000
1     1000
2     1000
3     1000
4     1000
....
44    1000
45    1000
46    1000
47    1000
48    1000
49    1000
Name: text, dtype: int64

因此,您会为每个 row 条目获得 .count():

txt = txt.apply(lambda x: nltk.Text(x).count('abac'))
txt.head()

0    27
1    24
2    17
3    25
4    32

然后您可以使用以下方法对结果求和:

txt.sum()

1239

【讨论】:

  • 感谢@Stefan,这几乎解决了我的问题,但是txt 对象仍然是熊猫数据框对象,这意味着我只能使用apply、map 或for 循环。但是,如果我想做nltk.Text(txt).concordance("the") 之类的事情,我会遇到问题。为了解决这个问题,我仍然需要将整个文本变量转换为字符串,正如我们在第一个示例中看到的那样,该字符串将因某种原因被截断。关于如何克服这个问题的任何想法?非常感谢!
  • 您可以将整个text column 转换为一个单词列表,使用:[t for t in df.text.tolist()] - 在创建之后或.tokenize() 之后。
猜你喜欢
  • 2012-09-17
  • 2017-01-20
  • 2017-06-18
  • 2019-03-14
  • 2020-06-21
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多