【发布时间】:2020-09-15 05:05:25
【问题描述】:
拆分训练和测试数据后,我想在熊猫上使用 sklearn TFIdfVectorizer
这是拆分数据的代码:
train = data_df
train_df,test_df= train_test_split(train,test_size=0.2)
我尝试使用 TFIdfVectorizer 函数:
start = time.clock()
vect = CountVectorizer(ngram_range=(2,2))
train_df = vect.fit_transform(train_df)
test_df = vect.transform(test_df)
print (time.clock()-start)
但它出现了这样的错误:
ValueError Traceback (most recent call last)
<ipython-input-36-3588531e9fc6> in <module>
3 vect = CountVectorizer(ngram_range=(2,2))
4 #converting traning features into numeric vector
----> 5 train_df = vect.fit_transform(train_df)
6 #converting training labels into numeric vector
7 test_df = vect.transform(test_df)
~\anaconda3\lib\site-packages\sklearn\feature_extraction\text.py in fit_transform(self, raw_documents, y)
1218
1219 vocabulary, X = self._count_vocab(raw_documents,
-> 1220 self.fixed_vocabulary_)
1221
1222 if self.binary:
~\anaconda3\lib\site-packages\sklearn\feature_extraction\text.py in _count_vocab(self, raw_documents, fixed_vocab)
1148 vocabulary = dict(vocabulary)
1149 if not vocabulary:
-> 1150 raise ValueError("empty vocabulary; perhaps the documents only"
1151 " contain stop words")
1152
ValueError: empty vocabulary; perhaps the documents only contain stop words
有什么我想念的吗?或解决此问题的任何解决方案?谢谢
【问题讨论】:
标签: python pandas machine-learning scikit-learn