【发布时间】:2014-03-31 15:56:49
【问题描述】:
我正在使用正则表达式和漂亮的汤从一篇文章中获取信息。我目前似乎无法从输出中得到我真正需要的东西。对于日期,我只需要获取列表中返回的第一个实例。我尝试遍历列表,但还没有多少运气。对于作者,我想去掉 a href 标签,只知道它是谁,而不是整个返回的字符串。我尝试了一个循环并更改了一些正则表达式调用,但无法缩小范围。任何指导将不胜感激。下面是相关代码:
import urllib2
from bs4 import BeautifulSoup
import re
from time import *
url: http://www.reuters.com/article/2014/02/26/us-afghanistan-usa-militants-idUSBREA1O1SV20140226
# Parse HTML of article, aka making soup
soup = BeautifulSoup(urllib2.urlopen(url).read())
# Write the article author to the file
regex = '<p class="byline">(.+?)</p>'
pattern = re.compile(regex)
byline = re.findall(pattern,str(soup))
txt.write("Author: " + str(byline) + '\n' + '\n')
# Write the article date to the file
regex = '<span class="timestamp">(.+?)</span>'
pattern = re.compile(regex)
byline = re.findall(pattern,str(soup))
txt.write("Date: " + str(byline) + '\n' + '\n')
【问题讨论】:
-
你根本不需要正则表达式,使用 BeautifulSoup!日期在网址的最后 8 个字符中。
-
你能提供一个例子来说明如何使用 bs4 抓住作者吗?我已经阅读了漂亮的汤文档,他们的方法没有产生所需的输出。虽然我是 python 新手,所以这可能是我的误解。
标签: python regex web-scraping