【发布时间】:2019-04-01 16:46:21
【问题描述】:
我想做 n-gram 方法,但是一个字母一个字母
普通 N-gram:
sentence : He want to watch football match
result:
he, he want, want, want to , to , to watch , watch , watch football , football, football match, match
我想这样做,但一个字母一个字母:
word : Angela
result:
a, an, n , ng , g , ge, e ,el, l , la ,a
这是我使用 Sklearn 的代码,但它仍然是逐字逐句而不是逐字母:
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer(ngram_range=(1, 100),token_pattern = r"(?u)\b\w+\b")
corpus = ['Angel','Angelica','John','Johnson']
X = vectorizer.fit_transform(corpus)
analyze = vectorizer.build_analyzer()
print(vectorizer.get_feature_names())
print(vectorizer.transform(['Angela']).toarray())
【问题讨论】:
标签: python scikit-learn nlp