【问题标题】:Python Requests ScrapingPython 请求抓取
【发布时间】:2014-03-04 18:04:19
【问题描述】:

我有一个要抓取的 URL 列表。

如果我单独使用每个 URL,我的代码在列表中有效;但是,当我将 URLS 存储在一个文件中并在循环中使用它时,它只会上升到第二个 URL 并停在第三个。

这是我的代码:

urls=open("file.txt")

url=urls.read()


main=url.split("\n")

url_number=0
while url_number<len(main):
    page = requests.get(main[url_number])
    tree = html.fromstring(page.text)


    tournament = tree.xpath('//title/text()')

    round1= tree.xpath('//div[@data-round]/span/text()')
    scoreup= tree.xpath('//div[contains(@class, "top_score")]/text()')
    scoredown= tree.xpath('//div[contains(@class, "bottom_score")]/text()')

    url_number=url_number+1
    print url_number

    print "\n"
    results = [] 
    score_number=0
    round_number=0
    match_number=0

    while round_number < len(round1):
        match_number +=1
        results.append(
                        [match_number,
                        round1[round_number],
                        scoreup[score_number],
                        round1[round_number+1],
                        scoredown[score_number],
                        tournament,])
        round_number=round_number+2
        score_number=score_number+1


    print results

此代码为我提供了第二个 URL,并且只为第三个 URL 打印 3,即 url_number,后跟此错误。

scoredown[score_number],
IndexError: list index out of range

【问题讨论】:

  • 您显示的错误清楚地表明问题出在scoredown[score_number] 行。显然,您的抓取逻辑不适用于该特定 URL。您应该手动检查该页面并调整您的逻辑。
  • 但是当我直接在代码中使用该网址时,它可以正常工作
  • @bluenic 错误总是发生在第三个 URL 上,不管是哪个 URL,也不管总共有多少个?

标签: python python-requests


【解决方案1】:

我已经重写了它以利用 Python 迭代器而不是索引循环:

from itertools import count, cycle, islice, izip
import lxml.html as lh
import requests

URL_FILE   = "file.txt"
ROW_FORMAT = "{:>5}     {:18} {:>5}     {:18} {:>5}     {}".format
HEADER     = ['Match', 'Up', 'Score', 'Down', 'Score', 'Tournament']

lookfor = {
    'tournament': '//title/text()',
    'round1':     '//div[@data-round]/span/text()',
    'scoreup':    '//div[contains(@class, "top_score")]/text()',
    'scoredown':  '//div[contains(@class, "bottom_score")]/text()'
}

def main():
    with open(URL_FILE) as inf:
        urls = (line.strip() for line in inf)
        for num,url in enumerate(urls, 1):
            # get data
            txt = requests.get(url).text
            tree = lh.fromstring(txt)

            # pull out the bits we want        
            scraped = {name:tree.xpath(path) for name,path in lookfor.items()}
            # (and fix the title)
            title = scraped['tournament']
            tournament = title[0].replace('\n', '') if title else ''

            # reslice the data so it lines up        
            match_num   = count(1)                               # 1, 2, 3, ...
            up_rounds   = islice(scraped['round1'], 0, None, 2)  # even rounds
            down_rounds = islice(scraped['round1'], 1, None, 2)  # odd rounds
            tournament  = cycle([tournament])      # repeats the name
            # ... and reassemble it        
            results = izip(match_num, up_rounds, scraped['scoreup'], down_rounds, scraped['scoredown'], tournament)

            # generate output
            print("\nRound {}:".format(num))

            print(ROW_FORMAT(*HEADER))
            for row in results:
                print(ROW_FORMAT(*row))

if __name__=="__main__":
    main()

在给定的 url 上,结果是:

Round 1:
Match     Up                 Score     Down               Score     Tournament
    1     Halcyon.680            2     Dubrick.528            0     IG Spring 2011 Omega Divisional #11 - Challonge
    2     rabidsnowman.208       0     Drunkenboi.856         2     IG Spring 2011 Omega Divisional #11 - Challonge
    3     GoSuRum.612            2     Halcyon.680            1     IG Spring 2011 Omega Divisional #11 - Challonge
    4     Hummingbird.656        1     hammy.161              2     IG Spring 2011 Omega Divisional #11 - Challonge
    5     Cryptic.528            2     Divination.275         0     IG Spring 2011 Omega Divisional #11 - Challonge
    6     tqyrusecc.243          2     Kodak.775              1     IG Spring 2011 Omega Divisional #11 - Challonge
    7     coLrsvp.138            0     Drunkenboi.856         1     IG Spring 2011 Omega Divisional #11 - Challonge
    8     Sharo.803              0     ices.813               2     IG Spring 2011 Omega Divisional #11 - Challonge
    9     vpchance.970           2     hayes.848              0     IG Spring 2011 Omega Divisional #11 - Challonge
   10     Leif.812               0     Amalaxinaoum.405       0     IG Spring 2011 Omega Divisional #11 - Challonge
   11     GoSuRum.612            0     hammy.161              2     IG Spring 2011 Omega Divisional #11 - Challonge
   12     Cryptic.528            0     tqyrusecc.243          2     IG Spring 2011 Omega Divisional #11 - Challonge
   13     Drunkenboi.856         1     ices.813               2     IG Spring 2011 Omega Divisional #11 - Challonge
   14     vpchance.970           2     Amalaxinaoum.405       0     IG Spring 2011 Omega Divisional #11 - Challonge
   15     hammy.161              0     tqyrusecc.243          2     IG Spring 2011 Omega Divisional #11 - Challonge
   16     ices.813               1     vpchance.970           2     IG Spring 2011 Omega Divisional #11 - Challonge
   17     tqyrusecc.243          0     vpchance.970           3     IG Spring 2011 Omega Divisional #11 - Challonge

【讨论】:

    【解决方案2】:

    我对您的代码有一些建议,它们应该可以帮助您避免以后出现此类问题。主要是没有理由使用所有这些 while 循环。您会更适合使用 for 循环显式循环遍历您的列表。

    此外,python 有很多内置函数,这意味着您不必做这些脏活。这两者一起可以将您的代码的第一部分转换为:

    for url in open('file.txt').readlines():
    

    如果不查看您正在抓取的网址,很难完全确定,但我敢打赌,您的退货列表的大小并不像您想象的那么一致。

    您的xpath 选择器看起来不是特别窄,因为如果您在round1、scoreup 或@ 中获得多个值,那么您使用的是while 循环而不是显式循环遍历结果987654326@ 列表,与您的预期略有不同,您会收到这样的错误。

    【讨论】:

    • 我还是 python 新手,这是 challonge.com/IG2011Springomega11 网站之一的示例
    • .readlines() 是多余的,因为文件句柄本身就是文件中行的迭代器。因此你可以这样做:for url in open('file.txt'): .
    猜你喜欢
    • 2022-01-18
    • 2021-03-24
    • 2018-03-22
    • 2021-10-28
    • 2016-02-21
    • 1970-01-01
    • 2021-12-27
    • 2021-09-05
    • 2020-04-20
    相关资源
    最近更新 更多