【问题标题】:How to use vectorized document output by Mahout seq2sparse in Apache Spark如何在 Apache Spark 中使用 Mahout seq2sparse 的矢量化文档输出
【发布时间】:2015-03-25 21:47:10
【问题描述】:

Mahout seq2sparse 生成一堆序列文件,如 here 的完整描述。我想使用具有这种格式的矢量化文档: (docID, TF-IDF Vector) 并从他的 TF-IDF 向量中创建一个JavaRDD<Vector>。有人可以指导我吗?

【问题讨论】:

    标签: hadoop apache-spark mahout


    【解决方案1】:

    此信息随时可用in the documentation。

    预处理:对于 PATH_TO_SEQUENCE_FILES 中的一组序列文件格式的文档,mahout seq2sparse 命令执行 TF-IDF 转换(-wt tfidf 选项)和 L2 长度归一化 (-n 2 选项)如下:

    $ mahout seq2sparse 
      -i ${PATH_TO_SEQUENCE_FILES} 
      -o ${PATH_TO_TFIDF_VECTORS} 
      -nv 
      -n 2
      -wt tfidf
    
    Training: The model is then trained using mahout spark-trainnb. The default is to train a Bayes model. The -c option is given to train
    

    C Bayes 模型:

    $ mahout spark-trainnb
      -i ${PATH_TO_TFIDF_VECTORS} 
      -o ${PATH_TO_MODEL}
      -ow 
      -c
    
    Label Assignment/Testing: Classification and testing on a holdout set can then be performed via mahout spark-testnb. Again, the -c
    

    option表示模型为Cbayes:

    $ mahout spark-testnb 
      -i ${PATH_TO_TFIDF_TEST_VECTORS}
      -m ${PATH_TO_MODEL} 
      -ow 
      -c
    

    查看mahout command script,我们看到它实际上使用了org.apache.mahout.drivers.TrainNBDriver 类。我们对<Text, VectorWritable> 类型的parts using 和TFIDF 向量感兴趣:

      /** Read the training set from inputPath/part-x-00000 sequence file of form <Text,VectorWritable> */
      private def readTrainingSet: DrmLike[_]= {
        val inputPath = parser.opts("input").asInstanceOf[String]
        val trainingSet= drm.drmDfsRead(inputPath)
        trainingSet
      }
    
      override def process(): Unit = {
        start()
    
        val complementary = parser.opts("trainComplementary").asInstanceOf[Boolean]
        val outputPath = parser.opts("output").asInstanceOf[String]
    
        val trainingSet = readTrainingSet
        val (labelIndex, aggregatedObservations) = SparkNaiveBayes.extractLabelsAndAggregateObservations(trainingSet)
        val model = NaiveBayes.train(aggregatedObservations, labelIndex)
    
        model.dfsWrite(outputPath)
    
        stop()
      }
    

    如果我们仔细观察,我们会看到输入正在被drm.drmDfsRead(inputPath) 调用转换。然后它将像这样转换(来自SparkEngine bindings 的示例)

      /**
       * Load DRM from hdfs (as in Mahout DRM format)
       *
       * @param path
       * @param sc spark context (wanted to make that implicit, doesn't work in current version of
       *           scala with the type bounds, sorry)
       *
       * @return DRM[Any] where Any is automatically translated to value type
       */
      def drmDfsRead (path: String, parMin:Int = 0)(implicit sc: DistributedContext): CheckpointedDrm[_] = {
    
        val drmMetadata = hdfsUtils.readDrmHeader(path)
        val k2vFunc = drmMetadata.keyW2ValFunc
    
        // Load RDD and convert all Writables to value types right away (due to reuse of writables in
        // Hadoop we must do it right after read operation).
        val rdd = sc.sequenceFile(path, classOf[Writable], classOf[VectorWritable], minPartitions = parMin)
    
            // Immediately convert keys and value writables into value types.
            .map { case (wKey, wVec) => k2vFunc(wKey) -> wVec.get()}
    
        // Wrap into a DRM type with correct matrix row key class tag evident.
        drmWrap(rdd = rdd, cacheHint = CacheHint.NONE)(drmMetadata.keyClassTag.asInstanceOf[ClassTag[Any]])
      }
    

    【讨论】:

    • 所以我使用 seq2sparse 将 txt 文档转换为您在帖子第一部分中提到的矢量化文档,但我对第二部分感到困惑。我应该使用哪部分代码将矢量化文档转换为 JavaRDD?此外,由于我的矢量化文档的大小很大,我需要在 Hadoop 中进行此转换。
    猜你喜欢
    • 2012-08-09
    • 1970-01-01
    • 2013-03-10
    • 2012-08-28
    • 2013-07-31
    • 2023-03-21
    • 2013-04-17
    • 2012-06-09
    • 1970-01-01
    相关资源
    最近更新 更多