【问题标题】:using regex to grab information from an article使用正则表达式从文章中获取信息
【发布时间】:2014-03-31 15:56:49
【问题描述】:

我正在使用正则表达式和漂亮的汤从一篇文章中获取信息。我目前似乎无法从输出中得到我真正需要的东西。对于日期,我只需要获取列表中返回的第一个实例。我尝试遍历列表,但还没有多少运气。对于作者,我想去掉 a href 标签,只知道它是谁,而不是整个返回的字符串。我尝试了一个循环并更改了一些正则表达式调用,但无法缩小范围。任何指导将不胜感激。下面是相关代码:

import urllib2
from bs4 import BeautifulSoup
import re
from time import *

url: http://www.reuters.com/article/2014/02/26/us-afghanistan-usa-militants-idUSBREA1O1SV20140226

# Parse HTML of article, aka making soup
soup = BeautifulSoup(urllib2.urlopen(url).read())

# Write the article author to the file    
regex = '<p class="byline">(.+?)</p>'
pattern = re.compile(regex)
byline = re.findall(pattern,str(soup))
txt.write("Author: " + str(byline) + '\n' + '\n')

# Write the article date to the file    
regex = '<span class="timestamp">(.+?)</span>'
pattern = re.compile(regex)
byline = re.findall(pattern,str(soup))
txt.write("Date: " + str(byline) + '\n' + '\n')

【问题讨论】:

  • 你根本不需要正则表达式,使用 BeautifulSoup!日期在网址的最后 8 个字符中。
  • 你能提供一个例子来说明如何使用 bs4 抓住作者吗?我已经阅读了漂亮的汤文档,他们的方法没有产生所需的输出。虽然我是 python 新手,所以这可能是我的误解。

标签: python regex web-scraping


【解决方案1】:

您可以使用 BeautifulSoup 使用与您描述的几乎相同的方法来准确获取您需要的内容,而无需使用正则表达式。既然你知道你感兴趣的标签的特性,你可以直接使用bs4的find进行搜索

#make some soup
soup = BeautifulSoup(urllib2.urlopen(url).read())

#extract byline and date text from their respective tags
try:
    byline=soup.find("p", {'class':'byline'}).text
    date=soup.find("span", {'class':'timestamp'}).text
except:
    print 'byline missing!'

更新: 如果将整个内容包装在 try/except 结构中,则可以解决缺少署名的情况并定义一些应该发生的替代操作。

【讨论】:

  • 您先生是个伟人。我已经为此苦苦挣扎了一段时间。感谢您花时间帮助初学者。
  • 如果署名不存在,它会破坏代码。有没有办法检查陈述是否属实?还是让它什么都不返回并继续而不中断?
  • 效果很好!再次感谢。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2022-07-20
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多