【问题标题】:Pyspark multilabel text classificationPyspark 多标签文本分类
【发布时间】:2018-05-16 01:04:36
【问题描述】:

我正在尝试预测未知文本的标签。我的数据如下所示:

+-----------------+-----------+
|      label      |   text    |
+-----------------+-----------+
| [0, 1, 0, 1, 0] | blah blah |
| [1, 1, 0, 0, 0] | foo bar   |
+-----------------+-----------+

使用多标签二值化方法编码的第一列。 我的管道:

tokenizer = Tokenizer(inputCol="text", outputCol="words")
hashingTF = HashingTF(inputCol=tokenizer.getOutputCol(), outputCol="features")
lsvc = LinearSVC(maxIter=10, regParam=0.1)
ovr = OneVsRest(classifier=lsvc)

pipeline = Pipeline(stages=[tokenizer, hashingTF, ovr])

model = pipeline.fit(result)

当我运行这段代码时,我收到了这个错误:

ValueError: invalid literal for int() with base 10: '[1, 0, 1, 0, 1, 1, 1, 0, 0]'

有什么想法吗?

【问题讨论】:

  • 你是如何生成标签列的。我正在做一些非常相似的事情,但我无法在 pyspark 中找到相当于 MultilableBinarizer 的东西。你能帮帮我吗?

标签: python apache-spark pyspark apache-spark-ml databricks


【解决方案1】:

查看错误

int() 的文字无效

我们看到问题在于标签的期望类型不是数组,而是对应样本类的单个值。换句话说,您需要将标签从多标签二值化编码转换为单个数字。

一种方法是首先将数组转换为字符串,然后使用StringIndexer:

to_string_udf = udf(lambda x: ''.join(str(e) for e in x), StringType())
df = df.withColumn("labelstring", to_string_udf(df.label))

indexer = StringIndexer(inputCol="labelstring", outputCol="label")
indexed = indexer.fit(df).transform(df)

这将为每个唯一数组创建一个单独的类别(类标签)。

【讨论】:

  • 这不会影响准确性吗?通过这种方法,我们预测标签语料库而不是单个标签
  • @KertisvanKertis:可能,但据我所知,如果您想在这里使用 PySpark,没有太多选择。如果数据集较小,您可以尝试sklearn,它确实具有支持此类标签的方法。
猜你喜欢
  • 2016-06-14
  • 2017-01-22
  • 2021-08-17
  • 2017-09-18
  • 2018-06-10
  • 2016-05-25
  • 2017-04-20
  • 2018-07-20
  • 2017-11-21
相关资源
最近更新 更多