【发布时间】:2018-07-31 08:45:27
【问题描述】:
我正在尝试用马达加斯加语(我的母语)创建一个带标签的语料库。我按照文档 Python text Processing 和 natural language processing 和页面 https://www.nltk.org/book/ch05.html 中的说明进行操作。 我已经设法基于通用词性标签集和一些带标签的语料库创建了自己的词性标签集。 这是我的代码:
import os, os.path
path = os.path.expanduser('D:/Mes documents/MY_POS_tagger/nltk_data')
if not os.path.exists(path):
os.mkdir(path)
print("OS path done :%s"%os.path.exists(path))
import nltk.data
nltk.data.path.append('D:/Mes documents/MY_POS_tagger/nltk_data')
print("NLTK data path done:%s"%(path in nltk.data.path))
#read a POSfile
import nltk
from nltk.corpus.reader import TaggedCorpusReader
from nltk.tag import UnigramTagger
#there's only one document malagasy.pos, it's there where my tagged corpora.
reader = TaggedCorpusReader('D:/Mes documents/MY_POS_tagger/nltk_data/corpora/cookbook', r'.*\.pos')
train_sents=reader.tagged_sents()
tagger=UnigramTagger(train_sents)
#dago.txt contain just sentences without tag, i just wanted to test if the tag i assign on the POS file will work
text=(nltk.data.load('corpora/cookbook/dago.txt', format='raw'))
text_tokenized=nltk.word_tokenize(text)
print tagger.tag(text_tokenized)
我有这个结果:
OS path done :True
NLTK data path done:True
[('Matory', u'VB'), ('ny', None), ('alika', u'NN')]
所以我可以看到它是有效的,但我在上面的文档中读到我必须训练我的标记器。所以我问是否有人可以建议我如何做到这一点,因为我读到我需要腌制一个训练有素的标记器并训练和组合 Ngram 标记器,但我不明白腌制的含义或作用。而且我不知道我现在正在做的是否是使用 NLTK 创建和利用标记语料库的正确路径。 谢谢
【问题讨论】:
-
这个链接可能对您有帮助吗? How to build POS-tagged corpus with NLTK?Veloma
-
谢谢你的建议,我会检查的:)
标签: python nltk corpus pos-tagger