【问题标题】:Regex parse Buffy Script using look behindsRegex 使用look behinds 解析 Buffy 脚本
【发布时间】:2017-04-29 22:50:14
【问题描述】:

我很难解析这个页面:http://www.buffyworld.com/buffy/transcripts/114_tran.html

我正在尝试获取带有相关对话的角色名称。 文本如下所示:

<p>BUFFY: Wait!
<p>She stands there panting, watching the truck turn a corner.
<p>BUFFY: (whining) Don't you want your garbage?
<p>She sighs, pouts, turns and walks back toward the house.
<p>Cut to the kitchen. Buffy enters through the back door, holding a pile of
mail. She begins looking through it. We see Dawn standing by the island.
<p>DAWN: Hey Buffy. Oh, don't forget, today's trash day.<br>BUFFY: (sourly)
Thanks.
<p>Dawn piles her books into her school bag. Buffy opens a letter.
<p>Close shot of the letter.
<p>
<p>Dawn smiles, and she and Willow exit. Buffy picks up the still-wrapped
sandwich and stares at it.
<p>BUFFY: (to herself) Somebody should.
<p>She sighs, puts the sandwich back in the bag.
<p>Cut to the Bronze. Pan across various people drinking and dancing,
bartender serving. Reveal Xander and Anya sitting at the bar eating chips from
several bags. A notebook sits in front of them bearing the wedding seating
chart.
<p>ANYA: See ... this seating chart makes no sense. We have to do it again.
(Xander nodding) We can't do it again. You do it.<br>XANDER: The seating
chart's fine. Let's get back to the table arrangements. I'm starting to have
dreams of gardenia bouquets. (winces) I am so glad my manly coworkers didn't
just hear me say that. (eating chips)

理想情况下,我会从&lt;p&gt; 或&lt;br&gt; 匹配到下一个&lt;p&gt; 或&lt;br&gt;。我试图为此使用前瞻后瞻:

reg = "((?<=<p>)|(?<=<br>))(?P<character>.+):(?P<dialogue>.+)((?=<p>)|(?=<br>))"
script = re.findall(reg, html_text)

很遗憾,这与任何内容都不匹配。当我离开前瞻((?=&lt;p&gt;)|(?=&lt;br&gt;)) 时,只要匹配对话中没有换行符,我就会匹配行。它似乎在换行符处终止,而不是继续到 &lt;p&gt;

例如。在这一行中,“谢谢”不匹配。 <p>DAWN: Hey Buffy. Oh, don't forget, today's trash day.<br>BUFFY: (sourly) Thanks.

感谢您提供的任何见解!

【问题讨论】:

  • 哪个蟒蛇?对我来说,第一个版本匹配的东西

标签: python regex regex-lookarounds


【解决方案1】:

解决点符号:

re.findall('((?<=<p>)|(?<=<br>))([A-Z]+):([^<]+)', text)

您也可以尝试使用special flag 将换行符包含在点的语义中。就个人而言,我什么时候可以使用拆分或一些 html 解析器。 RE 转义,所有参数、限制和标志都会让任何人发疯。还有re.split。

dialogs = {}
text = html_text.replace('<br>', '<p>')
paragraphs = text.split('<p>')

for p in paragraphs:
    if ":" in p:
        char, line = p.split(":", 1)
        if char in dialogs:
           dialogs[char].append(line)
        else:
           dialogs[char] = []

【讨论】:

  • 感谢您的帮助!
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-04-27
  • 1970-01-01
  • 1970-01-01
  • 2013-06-12
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多