【发布时间】:2018-05-16 01:04:36
【问题描述】:
我正在尝试预测未知文本的标签。我的数据如下所示:
+-----------------+-----------+
| label | text |
+-----------------+-----------+
| [0, 1, 0, 1, 0] | blah blah |
| [1, 1, 0, 0, 0] | foo bar |
+-----------------+-----------+
使用多标签二值化方法编码的第一列。 我的管道:
tokenizer = Tokenizer(inputCol="text", outputCol="words")
hashingTF = HashingTF(inputCol=tokenizer.getOutputCol(), outputCol="features")
lsvc = LinearSVC(maxIter=10, regParam=0.1)
ovr = OneVsRest(classifier=lsvc)
pipeline = Pipeline(stages=[tokenizer, hashingTF, ovr])
model = pipeline.fit(result)
当我运行这段代码时,我收到了这个错误:
ValueError: invalid literal for int() with base 10: '[1, 0, 1, 0, 1, 1, 1, 0, 0]'
有什么想法吗?
【问题讨论】:
-
你是如何生成标签列的。我正在做一些非常相似的事情,但我无法在 pyspark 中找到相当于 MultilableBinarizer 的东西。你能帮帮我吗?
标签: python apache-spark pyspark apache-spark-ml databricks