【问题标题】:POS tagging in ScalaScala中的POS标记
【发布时间】:2013-08-27 07:34:29
【问题描述】:

我尝试使用下面的斯坦福解析器在 Scala 中对句子进行 POS 标记

val lp:LexicalizedParser = LexicalizedParser.loadModel("edu/stanford/nlp/models/lexparser/englishPCFG.ser.gz");
lp.setOptionFlags("-maxLength", "50", "-retainTmpSubcategories")
val s = "I love to play"
val parse :Tree =  lp.apply(s)
val taggedWords = parse.taggedYield()
println(taggedWords)

我收到一个错误类型不匹配;在 val parse :Tree = lp.apply(s)

我不知道这是否是正确的做法。在 Scala 中还有其他简单的 POS 标记句子的方法吗?

【问题讨论】:

    标签: scala nlp stanford-nlp pos-tagger


    【解决方案1】:

    您不妨考虑使用 FACTORIE 工具包 (http://github.com/factorie/factorie)。它是一个用于机器学习和图形模型的通用库,恰好包含一套广泛的自然语言处理组件(标记化、标记规范化、形态分析、句子分割、词性标记、命名实体识别、依赖解析、提及发现,共指)。

    此外,它完全用 Scala 编写,并根据 Apache 许可证发布。

    文档目前很少,但会在未来几个月内得到改进。

    例如,基于 Maven 的安装完成后,您可以在命令行输入:

    bin/fac nlp --pos1 --parser1 --ner1
    

    启动一个套接字侦听的多线程 NLP 服务器。然后通过管道将纯文本传递到其套接字号来查询它:

    echo "Mr. Jones took a job at Google in New York.  He and his Australian wife moved from New South Wales on 4/1/12." | nc localhost 3228
    

    然后是输出

    1       1       Mr.             NNP     2       nn      O
    2       2       Jones           NNP     3       nsubj   U-PER
    3       3       took            VBD     0       root    O
    4       4       a               DT      5       det     O
    5       5       job             NN      3       dobj    O
    6       6       at              IN      3       prep    O
    7       7       Google          NNP     6       pobj    U-ORG
    8       8       in              IN      7       prep    O
    9       9       New             NNP     10      nn      B-LOC
    10      10      York            NNP     8       pobj    L-LOC
    11      11      .               .       3       punct   O
    
    12      1       He              PRP     6       nsubj   O
    13      2       and             CC      1       cc      O
    14      3       his             PRP$    5       poss    O
    15      4       Australian      JJ      5       amod    U-MISC
    16      5       wife            NN      6       nsubj   O
    17      6       moved           VBD     0       root    O
    18      7       from            IN      6       prep    O
    19      8       New             NNP     9       nn      B-LOC
    20      9       South           NNP     10      nn      I-LOC
    21      10      Wales           NNP     7       pobj    L-LOC
    22      11      on              IN      6       prep    O
    23      12      4/1/12          NNP     11      pobj    O
    24      13      .               .       6       punct   O
    

    当然,所有这些功能也有一个编程 API。

    import cc.factorie._
    import cc.factorie.app.nlp._
    val doc = new Document("Education is the most powerful weapon which you can use to change the world.")
    DocumentAnnotatorPipeline(pos.POS1).process(doc)
    for (token <- doc.tokens)
      println("%-10s %-5s".format(token.string, token.posLabel.categoryValue))
    

    将输出:

    Education  NN   
    is         VBZ  
    the        DT   
    most       RBS  
    powerful   JJ   
    weapon     NN   
    which      WDT  
    you        PRP  
    can        MD   
    use        VB   
    to         TO   
    change     VB   
    the        DT   
    world      NN   
    .          .    
    

    【讨论】:

    • 我刚刚在构建路径中添加了factorie-1.0.0-M6.jar,它在 DocumentAnnotatorPipeline.process(pos.POS1, doc) 上显示错误“未找到:值 DocumentAnnotatorPipeline”。我是否需要添加更多的罐子才能使其正常工作?
    • 您需要来自 GitHub 的最新版本。我们希望在几周内再创一个里程碑。
    【解决方案2】:

    我发现了一种在 Scala 中进行 POS 标记的非常简单的方法

    第 1 步

    从下面的链接下载 stanford tagger 3.2.0 版

    http://nlp.stanford.edu/software/stanford-postagger-2013-06-20.zip

    第 2 步

    将文件夹中的 stanford-postagger jar 添加到您的项目中,并将 english-left3words-distim.tagger 文件放在项目的模型文件夹中

    然后,使用下面的代码,您可以在 Scala 中对句子进行 pos 标记

                  val tagger = new MaxentTagger(
                    "english-left3words-distsim.tagger")
                  val art_con = "My name is Rahul"
                  val tagged = tagger.tagString(art_con)
                  println(tagged)
    

    输出: My_PRP$ name_NN is_VBZ Rahul_NNP

    【讨论】:

    • 有 scala 绑定使它更容易!安装需要一些工作,但随后整个事情被压缩成一行。图书馆是here
    【解决方案3】:

    我相信斯坦福解析器的 API 已经发生了一些变化,就像它有时一样。 apply 具有签名 public Tree apply(java.util.List&lt;? extends HasWord&gt; words),这就是您在错误消息中看到的内容。

    你现在应该使用的是parse,它的签名是public Tree parse(java.lang.String sentence)。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-10-30
      • 2016-07-02
      • 1970-01-01
      • 2012-06-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多