【发布时间】:2021-01-07 19:42:42
【问题描述】:
我有以下清单:
[
['the', 'the +Det'],
['dog', 'dog +N +A-right'],
['ran', 'run +V +past'],
['at', 'at +P'],
['me', 'I +N +G-left'],
['and', 'and +Cnj'],
['the', 'the +Det'],
['ball', 'ball +N +G-right'],
['was', 'was +C'],
['kicked', 'kick +V +past']
['by', 'by +P']
['me', 'I +N +A-left']
]
基本上,我想做的是:
- 遍历列表列表
- 查找
+G-left、+A-left、+G-right和+A-right的所有实例 - 如果看到
+G-left或+A-left,则向后查看带有元素+V的列表的第一个实例,将包含+G-left或+A-left的列表的第一个索引添加到末尾包含+V的列表中带有+G-left或+A-left标签,然后继续并重复 - 如果看到
+G-right或+A-right,期待第一个具有+V元素的列表实例,将包含+G-right或+A-right的列表的第一个索引添加到末尾包含+V的列表中带有+G-right或+A-right标签,然后继续并重复
所以在我上面的例子中,期望的状态是:
[
['the', 'the +Det'],
['dog', 'dog +N +A-right'],
['ran', 'run +V +past', 'dog+A-right', 'me+G-left'],
['at', 'at +P'],
['me', 'I +N +G-left'],
['and', 'and +Cnj'],
['the', 'the +Det'],
['ball', 'ball +N +G-right'],
['was', 'was +C'],
['kicked', 'kick +V +past', 'ball+G-right', 'me+A-left']
['by', 'by +P']
['me', 'I +N +A-left']
]
我认为解决这个问题的正确方法是使用re,所以:
gleft = re.compile(r"G-left")
gright = re.compile(r"G-right")
aleft = re.compile(r"A-left")
aright = re.compile(r"A-right")
然后类似
for item in list:
if aleft.match(item[1]):
somehow work backwards to find the +V tag
whatever.insert(-1, item[0]) #can you concatenate a string here to add +A-left
if aright.match(item[1]):
somehow work forwards to find the +V tag
whatever.insert(-1, item[0]) #can you concatenate a string here to add +A-right
同样的东西,但带有 G 标签。
希望有人可以帮助我指出正确的方向。我相信我已经正确地分解了这些步骤,我只是对 Python 不够熟悉,还不知道这件事的语法。
【问题讨论】:
-
我觉得周围所有这些符号真的让人头晕目眩。如果你能用一个更简单的例子来解释会不会很容易?
-
find all instances of正则表达式是:\+<@GR、\+<@AR、\+@GR>和\+@AR>但由于它们都是常量文字,因此您不需要正则表达式,使用 substr 或喜欢。 -
@Austin,我相信我已经让它变得更简单了。基本上我想把动词的主语和宾语放在动词的分析中。我正在研究的语言没有包,所以我不能使用预制树库。