【问题标题】:How to delete text within most outer loop如何删除最外循环中的文本
【发布时间】:2019-06-21 08:11:04
【问题描述】:

我正在尝试使用 python 中的正则表达式删除网页上最外层括号内包含链接的所有文本,但无济于事。

我尝试了一些正则表达式模式,如下所示:

paragraph = re.sub(r'\(.*[<a]+\)', '', p)

我试图检查标签是否存在于最外面的括号之间。

在这个来自维基百科的例子中:

    Rwanda (/ruˈɑːndə, -ˈæn-/ (About this soundlisten); Kinyarwanda: U Rwanda [u.ɾɡwaː.nda] (About this soundlisten)), officially the Republic of Rwanda (Kinyarwanda: Repubulika y'u Rwanda; Swahili: Jamhuri ya Rwanda; French: République du Rwanda) , is a country in Central ...

输入文字:

'<p><b>Rwanda</b> (<span class="nowrap"><span class="IPA nopopups noexcerpt"><a href="/wiki/Help:IPA/English" title="Help:IPA/English">/<span style="border-bottom:1px dotted"><span title="\'r\' in \'rye\'">r</span><span title="/u/: \'u\' in \'influence\'">u</span><span title="/ˈ/: primary stress follows">ˈ</span><span title="/ɑː/: \'a\' in \'father\'">ɑː</span><span title="\'n\' in \'nigh\'">n</span><span title="\'d\' in \'dye\'">d</span><span title="/ə/: \'a\' in \'about\'">ə</span></span>, <wbr/>-<span style="border-bottom:1px dotted"><span title="/ˈ/: primary stress follows">ˈ</span><span title="/æ/: \'a\' in \'bad\'">æ</span><span title="\'n\' in \'nigh\'">n</span></span>-/</a></span> <span class="nowrap" style="font-size:85%"><bracket><span class="unicode haudio"><span class="fn"><span style="white-space:nowrap;margin-right:.25em;"><a href="/wiki/File:Rwanda_pronunciation.ogg" title="About this sound"><img alt="About this sound" data-file-height="20" data-file-width="20" decoding="async" height="11" src="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/11px-Loudspeaker.svg.png" srcset="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/17px-Loudspeaker.svg.png 1.5x, //upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/22px-Loudspeaker.svg.png 2x" width="11"/></a></span><a class="internal" href="//upload.wikimedia.org/wikipedia/commons/2/2c/Rwanda_pronunciation.ogg" title="Rwanda pronunciation.ogg">listen</a></span></span>)</span></span>; <a class="mw-redirect" href="/wiki/Kinyarwanda_language" title="Kinyarwanda language">Kinyarwanda</a>:  <small></small><span class="IPA" title="Representation in the International Phonetic Alphabet <bracket>IPA)"><a href="/wiki/Help:IPA" title="Help:IPA">[u.ɾɡwaː.nda]</a></span> <span class="nowrap" style="font-size:85%"><bracket><span class="unicode haudio"><span class="fn"><span style="white-space:nowrap;margin-right:.25em;"><a href="/wiki/File:Rwanda_<bracket>rw)_pronunciation.ogg" title="About this sound"><img alt="About this sound" data-file-height="20" data-file-width="20" decoding="async" height="11" src="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/11px-Loudspeaker.svg.png" srcset="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/17px-Loudspeaker.svg.png 1.5x, //upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/22px-Loudspeaker.svg.png 2x" width="11"/></a></span><a class="internal" href="//upload.wikimedia.org/wikipedia/commons/3/34/Rwanda_%28rw%29_pronunciation.ogg" title="Rwanda <bracket>rw) pronunciation.ogg">listen</a></span></span>)</span>), officially the <b>Republic of Rwanda</b> <bracket><a class="mw-redirect" href="/wiki/Kinyarwanda_language" title="Kinyarwanda language">Kinyarwanda</a>: ; <a class="mw-redirect" href="/wiki/Kiswahili" title="Kiswahili">Swahili</a>: ; <a href="/wiki/French_language" title="French language">French</a>: ), is a country  in <a href="/wiki/Central_Africa" title="Central Africa">Central</a> ... </p>'

我希望输出如下:

Rwanda, officially the Republic of Rwanda ...

但是;它失败了,它从第一个左括号到最后一个左括号获取所有文本,而不是获取第一组外括号。

我可以使用正则表达式来做到这一点,还是必须寻找其他地方?

【问题讨论】:

  • 在处理网页时使用适当的 XML/HTML 解析器怎么样?因为那是正确的方式
  • 我正在做一个任务,我正在使用 python 解决哲学问题。
  • 请给出明确的输入 -> 想要的输出
  • 完成,请查看。
  • 对于最多 2 层嵌套(比您的样本多 1 层),请尝试 like this。

标签: python regex


【解决方案1】:

您可以使用 BeautifulSoup 将此问题转换为解析 HTML 页面之类的问题(假设括号是平衡的):

s = '''
Rwanda (/ruˈɑːndə, -ˈæn-/ (About this soundlisten); Kinyarwanda: U Rwanda [u.ɾɡwaː.nda] (About this soundlisten)), officially the Republic of Rwanda (Kinyarwanda: Repubulika y'u Rwanda; Swahili: Jamhuri ya Rwanda; French: République du Rwanda)
'''

import re
from bs4 import BeautifulSoup

s = re.sub(r'\(', r'<bracket>', s)
s = re.sub(r'\)', r'</bracket>', s)

soup = BeautifulSoup(s, 'lxml')
for bracket in soup.select('bracket'):
    bracket.extract()

s = re.sub(r'\s+,', r',', soup.body.text.strip())

print(s)

打印:

Rwanda, officially the Republic of Rwanda

【讨论】:

  • 这很好,但如果此文本包含链接,我的目标是删除括号内的所有文本,因为我想要第一个不在括号内的链接以供将来使用。当我将其转换回汤对象时,您的答案将成为主要标签中的所有文本并丢弃 标签。希望我有任何意义。
  • @amro_ghoneim 你能提供输入文本的例子吗?您可以编辑您的问题并将其放在那里。
  • 我刚做了,请检查一下。我想要的是删除括号内包含链接的所有文本,以便我可以获得括号外的第一个链接(在本例中为中非)。
猜你喜欢
  • 2013-01-10
  • 1970-01-01
  • 1970-01-01
  • 2013-12-20
  • 2018-11-03
  • 1970-01-01
  • 1970-01-01
  • 2018-01-22
  • 1970-01-01
相关资源
最近更新 更多