【问题标题】:How to tfidfvectorizer on pandas dataframe?如何在熊猫数据框上使用 tfidfvectorizer?
【发布时间】:2020-09-15 05:05:25
【问题描述】:

拆分训练和测试数据后,我想在熊猫上使用 sklearn TFIdfVectorizer

这是拆分数据的代码:

train = data_df
    train_df,test_df= train_test_split(train,test_size=0.2)

我尝试使用 TFIdfVectorizer 函数:

start = time.clock()
vect = CountVectorizer(ngram_range=(2,2))
train_df = vect.fit_transform(train_df)
test_df = vect.transform(test_df)

print (time.clock()-start)

但它出现了这样的错误:

ValueError                                Traceback (most recent call last)
<ipython-input-36-3588531e9fc6> in <module>
      3 vect = CountVectorizer(ngram_range=(2,2))
      4 #converting traning features into numeric vector
----> 5 train_df = vect.fit_transform(train_df)
      6 #converting training labels into numeric vector
      7 test_df = vect.transform(test_df)

~\anaconda3\lib\site-packages\sklearn\feature_extraction\text.py in fit_transform(self, raw_documents, y)
   1218 
   1219         vocabulary, X = self._count_vocab(raw_documents,
-> 1220                                           self.fixed_vocabulary_)
   1221 
   1222         if self.binary:

~\anaconda3\lib\site-packages\sklearn\feature_extraction\text.py in _count_vocab(self, raw_documents, fixed_vocab)
   1148             vocabulary = dict(vocabulary)
   1149             if not vocabulary:
-> 1150                 raise ValueError("empty vocabulary; perhaps the documents only"
   1151                                  " contain stop words")
   1152 

ValueError: empty vocabulary; perhaps the documents only contain stop words

有什么我想念的吗?或解决此问题的任何解决方案?谢谢

【问题讨论】:

    标签: python pandas machine-learning scikit-learn


    【解决方案1】:

    问题似乎是您的记录可能包含单个字符串。尝试将它们转换为列表或将最小文档频率设置为 1。请查看下面给出的链接,它会产生您想要的结果:

    ValueError: empty vocabulary; perhaps the documents only contain stop words

    【讨论】:

      猜你喜欢
      • 2020-02-17
      • 2018-12-25
      • 2019-05-14
      • 2019-03-21
      • 1970-01-01
      • 2019-07-28
      • 2022-11-13
      • 2015-10-23
      • 2019-01-19
      相关资源
      最近更新 更多