【问题标题】:What is the correct svmlight input format in Mallet?Mallet 中正确的 svmlight 输入格式是什么?
【发布时间】:2017-09-21 11:40:49
【问题描述】:

我使用Mallet 和SVMLight 输入格式来做classification 使用NaiveBayes 分类器。但我得到了NumberFormatException。我想知道在使用 SVMLight 时如何使用字符串功能。正如我在指南1 中所读到的,这些特征也可以是字符串。

谁能帮助我的代码或输入有什么问题?

这是我的代码:

public void trainMalletNaiveBayes() throws Exception {

        ArrayList<Pipe> pipes = new ArrayList<Pipe>();
        pipes.add(new SvmLight2FeatureVectorAndLabel());
        pipes.add(new PrintInputAndTarget());

        SerialPipes pipe = new SerialPipes(pipes);

        //prepare training instances
        InstanceList trainingInstanceList = new InstanceList(pipe);

        trainingInstanceList.addThruPipe(new CsvIterator(new FileReader("/tmp/featureFiles_svm.csv"), "^(\\S*)[\\s,]*(.*)$", 2, 1, -1));

        //prepare test instances
        InstanceList testingInstanceList = new InstanceList(pipe);
        testingInstanceList.addThruPipe(new CsvIterator(new FileReader("/tmp/test_set.csv"), "^(\\S*)[\\s,]*(.*)$", 2, 1, -1));

        ClassifierTrainer trainer = new NaiveBayesTrainer();
        Classifier classifier = trainer.train(trainingInstanceList);

这是我输入文件的前三行:

No f1:NP f2:NN f3:1 f4:1 f5:0 f6:0 f7:0 f8:0.0 f9:1 f10:true f11:false f12:false f13:false f14:false f15:ROOT f16:NN f17:NOTHING
No f1:NP f2:NN f3:8 f4:4 f5:0 f6:0 f7:1 f8:4.127134385045092 f9:8 f10:true f11:false f12:false f13:false f14:false f15:ROOT f16:DT f17:NOTHING
Yes f1:NP f2:NN f3:4 f4:3 f5:0 f6:0 f7:0 f8:0.0 f9:4 f10:true f11:false f12:false f13:false f14:false f15:NP f16:DT f17:NN

第一列是实例的标签,其余数据包括特征及其值。例如,NN 显示短语中心词的POS。

与此同时,我得到了 NN (NumberFormatException: For input string: "NN") 的例外情况。我想知道为什么之前的NP 没有任何问题,但停在NN。

【问题讨论】:

    标签: classification mallet svmlight


    【解决方案1】:

    所有特征都需要有数值。对于布尔值,您可以使用 true=1 和 false=0。您还必须将 f1:NP 修改为 f1_NP=1。

    它没有在 NP 上死掉的原因是 SvmLight2FeatureVectorAndLabel 类期望解析整行(标签和数据),但代码正在读取带有 CsvIterator 的文件,该文件正在拆分第一个元素作为标签。

    classify.tui.SvmLight2Vectors 类将此代码用于迭代器:

    new SelectiveFileLineIterator (fileReader, "^\\s*#.+")
    

    【讨论】:

    • 感谢您的回复。我是否应该将所有其他具有零值的功能添加到该行。例如,当我有一个特征的 NP 值时,这意味着它不是 VP、S、FRAG 等。我是否还要添加 f2_VP:0、f3_S:0 等?我的意思是,我是否应该转换我的分类数字特征的特征?然后,我将有一个非常稀疏的特征向量。对吗?
    • 将类别转换为特征,忽略任何零值,它将得到有效处理。
    • 谢谢。它现在可以正常工作了:) 只是另一个问题,使用上述格式和编写的代码,我得到名称:csvline:1 目标:f1_NP:1 输入:f2(0)=0.0 f3(1)=0.0 f4(2 )=2.65 ... 似乎它没有正确读取我的目标并将某个特征作为目标。我的代码或输入格式是否有问题,或者 PrintInputAndTarget() 不适用于 SVMLight,而仅适用于其他格式?
    • 哦,抱歉,我没有正确理解您的回答。我通过将 CSVIterator 更改为 SelectiveFileLineIterator 解决了我的问题,现在它可以完美运行了:)
    猜你喜欢
    • 1970-01-01
    • 2016-10-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-12-14
    • 2014-01-20
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多