【问题标题】:Fitting a model give error (ValueError: could not convert string to float:)拟合模型给出错误(ValueError:无法将字符串转换为浮点数:)
【发布时间】:2020-01-21 20:27:51
【问题描述】:

使用朴素贝叶斯算法

from sklearn.naive_bayes import MultinomialNB

nb = MultinomialNB()

代码一直工作到这一行,但是当我拟合模型时,它会显示错误。

nb.fit(X_train, y_train)

输出:

ValueError: could not convert string to float: 'My fiance and 
I tried the place because of a Groupon.  We live in the same neighborhood 
and see the place all the time but the look of the place was never enough 
to draw us in.  There is nothing eye catching about the business front at 
all.  It\'s in a strip mall and looks old..........

我正在使用 yelp.csv 数据集进行自然语言处理

预期的答案应该是这样的

MultinomialNB(alpha=1.0, class_prior=None, fit_prior=True)

【问题讨论】:

    标签: python machine-learning nlp training-data countvectorizer


    【解决方案1】:

    您收到此错误是因为 MultinomialNB 的 fit 方法需要浮点值的矩阵 X,而不是我猜你的文本提供。

    为了生成正确的 X 矩阵,您需要计算问题中每个类别的所有概率(例如,好或坏,您没有指定问题中的类别),你的词汇的概率分布,以及你的每类术语的概率。生成 X 矩阵后,您可以将其拟合到 MultinomialNB :)

    在这里您可以了解如何生成 X(请注意,这是有关构建多项式朴素贝叶斯分类器的完整教程,因此其中有很多内容您不需要你正在使用 sklearn):

    https://towardsdatascience.com/multinomial-naive-bayes-classifier-for-text-analysis-python-8dd6825ece67

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2019-09-28
      • 1970-01-01
      • 2019-03-23
      • 2018-06-13
      • 2013-05-30
      • 1970-01-01
      相关资源
      最近更新 更多