【问题标题】:How to use Python regular expression to get Image src?如何使用 Python 正则表达式获取 Image src?
【发布时间】:2012-06-13 21:14:06
【问题描述】:

如何使用正则表达式通过Python从下面的html字符串中获取图片的src

<td width="80" align="center" valign="top"><font style="font-size:85%;font-family:arial,sans-serif"><a href="http://news.google.com/news/url?sa=t&fd=R&usg=AFQjCNFqz8ZCIf6NjgPPiTd2LIrByKYLWA&url=http://www.news.com.au/business/spain-victory-faces-market-test/story-fn7mjon9-1226390697278"><img src="//nt3.ggpht.com/news/tbn/380jt5xHH6l_FM/6.jpg" alt="" border="1" width="80" height="80" /><br /><font size="-2">NEWS.com.au</font></a></font></td>

我尝试使用

matches = re.search('@src="([^"]+)"',text)
print(matches[0])

但一无所获

【问题讨论】:

标签: python html regex html-parsing


【解决方案1】:

您可以考虑使用BeautifulSoup,而不是正则表达式:

>>> from bs4 import BeautifulSoup
>>> soup = BeautifulSoup(junk)
>>> soup.findAll('img')
[<img src="//nt3.ggpht.com/news/tbn/380jt5xHH6l_FM/6.jpg" alt="" border="1" width="80" height="80" />]
>>> soup.findAll('img')[0]['src']
u'//nt3.ggpht.com/news/tbn/380jt5xHH6l_FM/6.jpg'

【讨论】:

  • Beautiful Soup 不会给解决方案增加很多开销吗? img 标签相对容易解析(而且由于它们不包含其他文本,通常格式正确)
【解决方案2】:

只需在正则表达式中丢失@,它就会起作用

【讨论】:

    【解决方案3】:

    您可以稍微简化您的re:

    match = re.search(r'src="(.*?)"', text)
    

    【讨论】:

    • 它也获取 javascript 文件。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-06-03
    相关资源
    最近更新 更多