【问题标题】:iterating for a beautifulsoup resultset python迭代一个beautifulsoup结果集python
【发布时间】:2015-04-08 01:27:33
【问题描述】:

我正在从网站 (http://sports.yahoo.com/nfl/players/8800/) 上抓取数据,为此我正在使用 urllib2 和 BeautifulSoup。我目前的代码如下所示:

site=  'http://sports.yahoo.com/nfl/players/8800/'
response = urllib2.urlopen(site)
html = response.read()
soup = BeautifulSoup(html)
rushing=[]
passing=[]
receiving=[]

#here is where my problem arises
for elem in soup.find_all('th', text=re.compile('2008')):
        passing = elem.parent.find_all('td', class_=re.compile('10'))
        rushing = elem.parent.find_all('td', class_=re.compile('20'))
        receiving = elem.parent.find_all('td', class_=re.compile('30'))

此页面上存在三种情况下,soup.find_all(...'2008')) 部分存在,并且在单独打印该部分时会出现这些情况。但是,运行这个 for 循环只会运行一次循环。如何确保循环运行 3 次?

【问题讨论】:

    标签: python html python-2.7 beautifulsoup html-parsing


    【解决方案1】:

    据我了解,您需要 extend() 在循环之前定义的列表:

    rushing = []
    passing = []
    receiving = []
    
    for elem in soup.find_all('th', text=re.compile('2008')):
        passing.extend([td.text for td in elem.parent.find_all('td', class_=re.compile('10'))])
        rushing.extend([td.text for td in elem.parent.find_all('td', class_=re.compile('20'))])
        receiving.extend([td.text for td in elem.parent.find_all('td', class_=re.compile('30'))])
    
    print passing
    print rushing
    print receiving
    

    打印:

    [u'3']
    [u'19', u'58', u'14.5', u'3.1', u'0']
    [u'2', u'17', u'4.3', u'8.5', u'11', u'6.5', u'0']
    

    【讨论】:

      猜你喜欢
      • 2020-01-02
      • 1970-01-01
      • 1970-01-01
      • 2014-10-02
      • 2014-10-25
      • 2017-03-07
      • 2021-09-29
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多