【问题标题】:Python locate words in nltk.treePython 在 nltk.tree 中定位单词
【发布时间】:2016-06-06 14:42:39
【问题描述】:

我正在尝试构建一个 nltk 来获取单词的上下文。我有两句话

sentences=pd.DataFrame({"sentence": ["The weather was good so I went swimming", "Because of the good food we took desert"]})

我想知道,“好”这个词指的是什么。我的想法是对句子进行分块(代码来自教程here),然后查看单词“good”和名词是否在同一个节点中。如果不是,它指的是在那之前或之后的名词。

首先我按照教程构建 Chunker

from nltk.corpus import conll2000
test_sents = conll2000.chunked_sents('test.txt', chunk_types=['NP'])
train_sents = conll2000.chunked_sents('train.txt', chunk_types=['NP'])

class ChunkParser(nltk.ChunkParserI):
    def __init__(self, train_sents):
        train_data = [[(t,c) for w,t,c in nltk.chunk.tree2conlltags(sent)]
            for sent in train_sents]
        self.tagger = nltk.TrigramTagger(train_data)
    def parse(self, sentence):
        pos_tags = [pos for (word,pos) in sentence]
        tagged_pos_tags = self.tagger.tag(pos_tags)
        chunktags = [chunktag for (pos, chunktag) in tagged_pos_tags]
        conlltags = [(word, pos, chunktag) for ((word,pos),chunktag)
        in zip(sentence, chunktags)]
        return nltk.chunk.conlltags2tree(conlltags)

NPChunker = ChunkParser(train_sents)

然后,我将其应用于我的句子:

sentence=sentences["sentence"][0]
tags=nltk.pos_tag(sentence.lower().split())
result = NPChunker.parse(tags)
print result

结果是这样的

(S
  (NP the/DT weather/NN)
  was/VBD
  (NP good/JJ)
  so/RB
  (NP i/JJ)
  went/VBD
  swimming/VBG)

现在我想“找到”单词“good”在哪个节点中。我还没有真正想出更好的方法,而是计算节点和叶子中的单词。 “好”这个词是句子中的第 3 个词。

stuctured_sentence=[]
for n in range(len(result)):
    stuctured_sentence.append(list(result[n]))

structure_length=[]
for n in result:
    if isinstance(n, nltk.tree.Tree):               
        if n.label() == 'NP':
            print n
            structure_length.append(len(n))
    else:
        print str(n) +"is a leaf"
        structure_length.append(1)

通过总结字数,我知道“好”这个词在哪里。

structure_frame=pd.DataFrame({"structure": stuctured_sentence, "length": structure_length})
structure_frame["cumsum"]=structure_frame["length"].cumsum()

有没有更简单的方法来确定单词的节点或叶子,并找出“好”指的是哪个单词?

最好的亚历克斯

【问题讨论】:

    标签: python nltk chunking


    【解决方案1】:

    在叶子列表中找到您的单词是最容易的。然后,您可以将叶索引转换为树索引,即树下的路径。要查看与 good 分组的内容,请上一层并检查它挑选出的子树。

    首先,找出good在你的平句中的位置。 (如果您仍然将未标记的句子作为标记列表,则可以跳过此步骤。)

    words = [ w for w, t in result.leaves() ]
    

    现在我们找到good的线性位置,并转化​​为树形路径:

    >>> position = words.index("good")
    >>> treeposition = result.leaf_treeposition(position)
    >>> print(treeposition)
    (2, 0)
    

    “树位置”是树下的路径,表示为元组。 (NLTK 树可以使用元组和整数进行索引。)要查看good 的姐妹,请在到达路径尽头之前停下一步。

    >>> print(result[ treeposition[:-1] ])
    Tree('NP', [('good', 'JJ')])
    

    你来了。一个有一个叶子的子树,一对(good, JJ)。

    【讨论】:

    • 谢谢!这帮助了我。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-08-12
    • 1970-01-01
    • 2017-09-27
    • 2015-09-15
    • 1970-01-01
    • 2023-01-27
    相关资源
    最近更新 更多