【问题标题】:user defined feature in CRF++CRF++ 中的用户定义特征
【发布时间】:2015-02-03 12:03:27
【问题描述】:

我尝试向 CRF++ 模板添加更多功能。

根据How can I tell CRF++ classifier that a word x is captilized or understanding punctuations?

训练样本

The  DT  0  1   0   1   B-MISC
Oxford  NNP 0   1   0   1   I-MISC
Companion   NNP 0   1   0   1   I-MISC
to  TO  0   0   0   0   I-MISC
Philosophy  NNP 0   1   0   1   I-MISC

功能模板

# Unigram
U00:%x[-2,0]
U01:%x[-1,0]
U02:%x[0,0]
U03:%x[1,0]
U04:%x[2,0]
U05:%x[-1,0]/%x[0,0]
U06:%x[0,0]/%x[1,0]
U07:%x[-2,0]/%x[-1,0]/%x[0,0]

#shape feature
U08:%x[-2,2]
U09:%x[-1,2]
U10:%x[0,2]
U11:%x[1,2]
U12:%x[2,2]

B

训练阶段没问题。但是我没有得到 crf_test 的输出

tilney@ubuntu:/data/wikipedia/en$ crf_test -m validation_model test.data
tilney@ubuntu:/data/wikipedia/en$ 

如果忽略上面的形状,一切正常。我哪里做错了?

【问题讨论】:

    标签: nlp crf crf++


    【解决方案1】:

    我想通了。这是我的测试数据的问题。我认为每个特征都应该取自训练好的模型,所以我的测试数据中只有两列:word tag,结果证明测试文件的格式应该与训练数据完全相同!

    【讨论】:

      猜你喜欢
      • 2010-12-14
      • 2016-07-31
      • 2014-11-26
      • 2022-01-24
      • 2021-08-28
      • 1970-01-01
      • 2017-11-09
      • 1970-01-01
      • 2021-02-24
      相关资源
      最近更新 更多