【发布时间】:2015-04-08 01:27:33
【问题描述】:
我正在从网站 (http://sports.yahoo.com/nfl/players/8800/) 上抓取数据,为此我正在使用 urllib2 和 BeautifulSoup。我目前的代码如下所示:
site= 'http://sports.yahoo.com/nfl/players/8800/'
response = urllib2.urlopen(site)
html = response.read()
soup = BeautifulSoup(html)
rushing=[]
passing=[]
receiving=[]
#here is where my problem arises
for elem in soup.find_all('th', text=re.compile('2008')):
passing = elem.parent.find_all('td', class_=re.compile('10'))
rushing = elem.parent.find_all('td', class_=re.compile('20'))
receiving = elem.parent.find_all('td', class_=re.compile('30'))
此页面上存在三种情况下,soup.find_all(...'2008')) 部分存在,并且在单独打印该部分时会出现这些情况。但是,运行这个 for 循环只会运行一次循环。如何确保循环运行 3 次?
【问题讨论】:
标签: python html python-2.7 beautifulsoup html-parsing