【问题标题】:Find first match backwards in a list of lists在列表列表中向后查找第一个匹配项
【发布时间】:2021-01-07 19:42:42
【问题描述】:

我有以下清单:

[
['the', 'the +Det'],
['dog', 'dog +N +A-right'],
['ran', 'run +V +past'],
['at', 'at +P'], 
['me', 'I +N +G-left'],
['and', 'and +Cnj'],
['the', 'the +Det'],
['ball', 'ball +N +G-right'],
['was', 'was +C'],
['kicked', 'kick +V +past']
['by', 'by +P']
['me', 'I +N +A-left']

]

基本上,我想做的是:

  1. 遍历列表列表
  2. 查找+G-left、+A-left、+G-right 和+A-right 的所有实例
  3. 如果看到+G-left 或+A-left,则向后查看带有元素+V 的列表的第一个实例,将包含+G-left 或+A-left 的列表的第一个索引添加到末尾包含+V 的列表中带有+G-left 或+A-left 标签,然后继续并重复
  4. 如果看到+G-right或+A-right,期待第一个具有+V元素的列表实例,将包含+G-right或+A-right的列表的第一个索引添加到末尾包含+V 的列表中带有+G-right 或+A-right 标签,然后继续并重复

所以在我上面的例子中,期望的状态是:

[
['the', 'the +Det'],
['dog', 'dog +N +A-right'],
['ran', 'run +V +past', 'dog+A-right', 'me+G-left'],
['at', 'at +P'], 
['me', 'I +N +G-left'],
['and', 'and +Cnj'],
['the', 'the +Det'],
['ball', 'ball +N +G-right'],
['was', 'was +C'],
['kicked', 'kick +V +past', 'ball+G-right', 'me+A-left']
['by', 'by +P']
['me', 'I +N +A-left']
]

我认为解决这个问题的正确方法是使用re,所以:

gleft = re.compile(r"G-left")
gright = re.compile(r"G-right")
aleft = re.compile(r"A-left")
aright = re.compile(r"A-right")

然后类似

for item in list:
    if aleft.match(item[1]):
        somehow work backwards to find the +V tag
            whatever.insert(-1, item[0]) #can you concatenate a string here to add +A-left

    if aright.match(item[1]):
        somehow work forwards to find the +V tag
            whatever.insert(-1, item[0]) #can you concatenate a string here to add +A-right

同样的东西,但带有 G 标签。

希望有人可以帮助我指出正确的方向。我相信我已经正确地分解了这些步骤,我只是对 Python 不够熟悉,还不知道这件事的语法。

【问题讨论】:

  • 我觉得周围所有这些符号真的让人头晕目眩。如果你能用一个更简单的例子来解释会不会很容易?
  • find all instances of 正则表达式是:\+<@GR、\+<@AR、\+@GR> 和 \+@AR> 但由于它们都是常量文字,因此您不需要正则表达式,使用 substr 或喜欢。
  • @Austin,我相信我已经让它变得更简单了。基本上我想把动词的主语和宾语放在动词的分析中。我正在研究的语言没有包,所以我不能使用预制树库。

标签: python regex list nlp


【解决方案1】:

这可能可以通过使用辅助函数来简化,但除此之外,试试这个,它不需要正则表达式:

wls = [your list of lists, above, fixed (some commas are missing)]
for wl in wls:
    for w in wl:
        if '-right' in w:                        
            targ = wls.index(wl)            
            counter = 0
            for wt in (wls[targ+1:]):                               
                for t in wt:
                    if '+V' in t:
                        if counter<1:                            
                            wt.insert(len(wt),wl[0]+w.split(' ')[-1])
                        counter+=1

        if '-left' in w:            
            targ = wls.index(wl)            
            counter = 0
            revd = [item for item in reversed(wls[:targ])]
            for wt in revd:           
                for t in wt:
                    if '+V' in t:
                        if counter<1:
                            wt.insert(len(wt),wl[0]+w.split(' ')[-1])
                        counter+=1
           
wls

输出应该是你要找的。​​p>

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-07-18
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-10-15
    • 2015-04-16
    相关资源
    最近更新 更多