【发布时间】:2015-03-25 21:47:10
【问题描述】:
Mahout seq2sparse 生成一堆序列文件,如 here 的完整描述。我想使用具有这种格式的矢量化文档: (docID, TF-IDF Vector) 并从他的 TF-IDF 向量中创建一个JavaRDD<Vector>。有人可以指导我吗?
【问题讨论】:
标签: hadoop apache-spark mahout
Mahout seq2sparse 生成一堆序列文件,如 here 的完整描述。我想使用具有这种格式的矢量化文档: (docID, TF-IDF Vector) 并从他的 TF-IDF 向量中创建一个JavaRDD<Vector>。有人可以指导我吗?
【问题讨论】:
标签: hadoop apache-spark mahout
此信息随时可用in the documentation。
预处理:对于 PATH_TO_SEQUENCE_FILES 中的一组序列文件格式的文档,mahout seq2sparse 命令执行 TF-IDF 转换(-wt tfidf 选项)和 L2 长度归一化 (-n 2 选项)如下:
$ mahout seq2sparse -i ${PATH_TO_SEQUENCE_FILES} -o ${PATH_TO_TFIDF_VECTORS} -nv -n 2 -wt tfidf Training: The model is then trained using mahout spark-trainnb. The default is to train a Bayes model. The -c option is given to trainC Bayes 模型:
$ mahout spark-trainnb -i ${PATH_TO_TFIDF_VECTORS} -o ${PATH_TO_MODEL} -ow -c Label Assignment/Testing: Classification and testing on a holdout set can then be performed via mahout spark-testnb. Again, the -coption表示模型为Cbayes:
$ mahout spark-testnb -i ${PATH_TO_TFIDF_TEST_VECTORS} -m ${PATH_TO_MODEL} -ow -c
查看mahout command script,我们看到它实际上使用了org.apache.mahout.drivers.TrainNBDriver 类。我们对<Text, VectorWritable> 类型的parts using 和TFIDF 向量感兴趣:
/** Read the training set from inputPath/part-x-00000 sequence file of form <Text,VectorWritable> */
private def readTrainingSet: DrmLike[_]= {
val inputPath = parser.opts("input").asInstanceOf[String]
val trainingSet= drm.drmDfsRead(inputPath)
trainingSet
}
override def process(): Unit = {
start()
val complementary = parser.opts("trainComplementary").asInstanceOf[Boolean]
val outputPath = parser.opts("output").asInstanceOf[String]
val trainingSet = readTrainingSet
val (labelIndex, aggregatedObservations) = SparkNaiveBayes.extractLabelsAndAggregateObservations(trainingSet)
val model = NaiveBayes.train(aggregatedObservations, labelIndex)
model.dfsWrite(outputPath)
stop()
}
如果我们仔细观察,我们会看到输入正在被drm.drmDfsRead(inputPath) 调用转换。然后它将像这样转换(来自SparkEngine bindings 的示例)
/**
* Load DRM from hdfs (as in Mahout DRM format)
*
* @param path
* @param sc spark context (wanted to make that implicit, doesn't work in current version of
* scala with the type bounds, sorry)
*
* @return DRM[Any] where Any is automatically translated to value type
*/
def drmDfsRead (path: String, parMin:Int = 0)(implicit sc: DistributedContext): CheckpointedDrm[_] = {
val drmMetadata = hdfsUtils.readDrmHeader(path)
val k2vFunc = drmMetadata.keyW2ValFunc
// Load RDD and convert all Writables to value types right away (due to reuse of writables in
// Hadoop we must do it right after read operation).
val rdd = sc.sequenceFile(path, classOf[Writable], classOf[VectorWritable], minPartitions = parMin)
// Immediately convert keys and value writables into value types.
.map { case (wKey, wVec) => k2vFunc(wKey) -> wVec.get()}
// Wrap into a DRM type with correct matrix row key class tag evident.
drmWrap(rdd = rdd, cacheHint = CacheHint.NONE)(drmMetadata.keyClassTag.asInstanceOf[ClassTag[Any]])
}
【讨论】:
seq2sparse 将 txt 文档转换为您在帖子第一部分中提到的矢量化文档,但我对第二部分感到困惑。我应该使用哪部分代码将矢量化文档转换为 JavaRDD