【发布时间】:2019-06-21 08:11:04
【问题描述】:
我正在尝试使用 python 中的正则表达式删除网页上最外层括号内包含链接的所有文本,但无济于事。
我尝试了一些正则表达式模式,如下所示:
paragraph = re.sub(r'\(.*[<a]+\)', '', p)
我试图检查标签是否存在于最外面的括号之间。
在这个来自维基百科的例子中:
Rwanda (/ruˈɑːndə, -ˈæn-/ (About this soundlisten); Kinyarwanda: U Rwanda [u.ɾɡwaː.nda] (About this soundlisten)), officially the Republic of Rwanda (Kinyarwanda: Repubulika y'u Rwanda; Swahili: Jamhuri ya Rwanda; French: République du Rwanda) , is a country in Central ...
输入文字:
'<p><b>Rwanda</b> (<span class="nowrap"><span class="IPA nopopups noexcerpt"><a href="/wiki/Help:IPA/English" title="Help:IPA/English">/<span style="border-bottom:1px dotted"><span title="\'r\' in \'rye\'">r</span><span title="/u/: \'u\' in \'influence\'">u</span><span title="/ˈ/: primary stress follows">ˈ</span><span title="/ɑː/: \'a\' in \'father\'">ɑː</span><span title="\'n\' in \'nigh\'">n</span><span title="\'d\' in \'dye\'">d</span><span title="/ə/: \'a\' in \'about\'">ə</span></span>, <wbr/>-<span style="border-bottom:1px dotted"><span title="/ˈ/: primary stress follows">ˈ</span><span title="/æ/: \'a\' in \'bad\'">æ</span><span title="\'n\' in \'nigh\'">n</span></span>-/</a></span> <span class="nowrap" style="font-size:85%"><bracket><span class="unicode haudio"><span class="fn"><span style="white-space:nowrap;margin-right:.25em;"><a href="/wiki/File:Rwanda_pronunciation.ogg" title="About this sound"><img alt="About this sound" data-file-height="20" data-file-width="20" decoding="async" height="11" src="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/11px-Loudspeaker.svg.png" srcset="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/17px-Loudspeaker.svg.png 1.5x, //upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/22px-Loudspeaker.svg.png 2x" width="11"/></a></span><a class="internal" href="//upload.wikimedia.org/wikipedia/commons/2/2c/Rwanda_pronunciation.ogg" title="Rwanda pronunciation.ogg">listen</a></span></span>)</span></span>; <a class="mw-redirect" href="/wiki/Kinyarwanda_language" title="Kinyarwanda language">Kinyarwanda</a>: <small></small><span class="IPA" title="Representation in the International Phonetic Alphabet <bracket>IPA)"><a href="/wiki/Help:IPA" title="Help:IPA">[u.ɾɡwaː.nda]</a></span> <span class="nowrap" style="font-size:85%"><bracket><span class="unicode haudio"><span class="fn"><span style="white-space:nowrap;margin-right:.25em;"><a href="/wiki/File:Rwanda_<bracket>rw)_pronunciation.ogg" title="About this sound"><img alt="About this sound" data-file-height="20" data-file-width="20" decoding="async" height="11" src="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/11px-Loudspeaker.svg.png" srcset="//upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/17px-Loudspeaker.svg.png 1.5x, //upload.wikimedia.org/wikipedia/commons/thumb/8/8a/Loudspeaker.svg/22px-Loudspeaker.svg.png 2x" width="11"/></a></span><a class="internal" href="//upload.wikimedia.org/wikipedia/commons/3/34/Rwanda_%28rw%29_pronunciation.ogg" title="Rwanda <bracket>rw) pronunciation.ogg">listen</a></span></span>)</span>), officially the <b>Republic of Rwanda</b> <bracket><a class="mw-redirect" href="/wiki/Kinyarwanda_language" title="Kinyarwanda language">Kinyarwanda</a>: ; <a class="mw-redirect" href="/wiki/Kiswahili" title="Kiswahili">Swahili</a>: ; <a href="/wiki/French_language" title="French language">French</a>: ), is a country in <a href="/wiki/Central_Africa" title="Central Africa">Central</a> ... </p>'
我希望输出如下:
Rwanda, officially the Republic of Rwanda ...
但是;它失败了,它从第一个左括号到最后一个左括号获取所有文本,而不是获取第一组外括号。
我可以使用正则表达式来做到这一点,还是必须寻找其他地方?
【问题讨论】:
-
在处理网页时使用适当的 XML/HTML 解析器怎么样?因为那是正确的方式
-
我正在做一个任务,我正在使用 python 解决哲学问题。
-
请给出明确的输入 -> 想要的输出
-
完成,请查看。
-
对于最多 2 层嵌套(比您的样本多 1 层),请尝试 like this。