使用来自inxi -C 的具有这些规格的服务器:
CPU(s): 2 Hexa core Intel Xeon CPU E5-2430 v2s (-HT-MCP-SMP-) cache: 30720 KB flags: (lm nx sse sse2 sse3 sse4_1 sse4_2 ssse3 vmx)
Clock Speeds: 1: 2500.036 MHz
通常,规范的答案是使用带有pos_tag_sents 的批量标记,但它似乎并不快。
让我们尝试在获得 POS 标签之前分析一些步骤(仅使用 1 个核心):
import time
from nltk.corpus import brown
from nltk import sent_tokenize, word_tokenize, pos_tag
from nltk import pos_tag_sents
# Load brown corpus
start = time.time()
brown_corpus = brown.raw()
loading_time = time.time() - start
print "Loading brown corpus took", loading_time
# Sentence tokenizing corpus
start = time.time()
brown_sents = sent_tokenize(brown_corpus)
sent_time = time.time() - start
print "Sentence tokenizing corpus took", sent_time
# Word tokenizing corpus
start = time.time()
brown_words = [word_tokenize(i) for i in brown_sents]
word_time = time.time() - start
print "Word tokenizing corpus took", word_time
# Loading, sent_tokenize, word_tokenize all together.
start = time.time()
brown_words = [word_tokenize(s) for s in sent_tokenize(brown.raw())]
tokenize_time = time.time() - start
print "Loading and tokenizing corpus took", tokenize_time
# POS tagging one sentence at a time took.
start = time.time()
brown_tagged = [pos_tag(word_tokenize(s)) for s in sent_tokenize(brown.raw())]
tagging_time = time.time() - start
print "Tagging sentence by sentence took", tagging_time
# Using batch_pos_tag.
start = time.time()
brown_tagged = pos_tag_sents([word_tokenize(s) for s in sent_tokenize(brown.raw())])
tagging_time = time.time() - start
print "Tagging sentences by batch took", tagging_time
[出]:
Loading brown corpus took 0.154870033264
Sentence tokenizing corpus took 3.77206301689
Word tokenizing corpus took 13.982845068
Loading and tokenizing corpus took 17.8847839832
Tagging sentence by sentence took 1114.65085101
Tagging sentences by batch took 1104.63432097
注意:pos_tag_sents 在 NLTK3.0 之前的版本中以前称为 batch_pos_tag
总之,我认为您需要考虑使用其他 POS 标记器来预处理您的数据,或者您必须使用 threading 来处理 POS 标记。