【问题标题】:Improve speed/performance of web-scraping with lots of exceptions提高网络抓取的速度/性能,但有很多例外
【发布时间】:2019-04-23 03:07:02
【问题描述】:

我编写了一些当前可以运行的网络抓取代码,但是速度很慢。一些背景:我正在使用 Selenium,因为它需要几个阶段的点击和进入,以及 BeautifulSoup。我的代码正在查看网站上子类别中的材料列表(下图)并抓取它们。如果从网站上抓取的材料是我感兴趣的 30 种材料之一(如下所示),那么它将数字 1 写入数据框,然后我将其转换为 Excel 工作表。

无论如何,我相信它之所以这么慢,是因为有很多例外。但是,除了try/except之外,我不确定如何处理这些。代码的主要部分如下所示,因为整段代码相当长。我还附上了相关网站的图片以供参考。

lst = ["Household cleaner and detergent bottles", "Plastic milk bottles", "Toiletries and shampoo bottles", "Plastic drinks bottles", 
       "Drinks cans", "Food tins", "Metal lids from glass jars", "Aerosols", 
       "Food pots and tubs", "Margarine tubs", "Plastic trays","Yoghurt pots", "Carrier bags",
       "Aluminium foil", "Foil trays",
       "Cardboard sleeves", "Cardboard egg boxes", "Cardboard fruit and veg punnets", "Cereal boxes", "Corrugated cardboard", "Toilet roll tubes", "Food and drink cartons",
       "Newspapers", "Window envelopes", "Magazines", "Junk mail", "Brown envelopes", "Shredded paper", "Yellow Pages" , "Telephone directories",
       "Glass bottles and jars"]

def site_scraper(site):
    page_loc = ('//*[@id="wrap-rlw"]/div/div[2]/div/div/div/div[2]/div/ol/li[{}]/div').format(site)
    page = driver.find_element_by_xpath(page_loc) 
    page.click()
    driver.execute_script("arguments[0].scrollIntoView(true);", page)

    soup=BeautifulSoup(driver.page_source, 'lxml')
    for i in x:
        for j in y:
            try:
                material = soup.find_all("div", class_ = "rlw-accordion-content")[i].find_all('li')[j].get_text(strip=True).encode('utf-8')
                if material in lst:
                    df.at[code_no, material] = 1 
                else:
                    continue 
                continue
            except IndexError:
                continue

x = xrange(0,8) 
y = xrange(0,9)

p = xrange(1,31)

for site in p:
    site_scraper(site)

具体来说,i 和 j 很少达到 6,7 或 8,但当它们出现时,我也必须捕获这些信息。对于上下文,i 对应于下图中不同类别的数量(汽车、建筑材料等),而 j 代表子列表(汽车电池和发动机油等)。因为每个代码的所有 30 个站点都重复这两个循环,而我有 1500 个代码,所以这非常慢。目前 10 个代码需要 6.5 分钟。

有没有办法改进这个过程?我尝试了列表理解,但是很难处理这样的错误,而且我的结果不再准确。 “如果”函数可能是一个更好的选择,如果是这样,我将如何合并它?如果需要,我也很乐意附上完整的代码。谢谢!

编辑: 通过改变

        except IndexError:
            continue

到

        except IndexError:
            break

它现在的运行速度几乎是原来的两倍!显然最好在失败一次后退出循环,因为后面的迭代也会失败。但是,仍然欢迎任何其他 pythonic 技巧:)

【问题讨论】:

    标签: python exception web-scraping error-handling try-catch


    【解决方案1】:

    听起来你只需要那些lis的文字:

    lis = driver.execute_script("[...document.querySelectorAll('.rlw-accordion-content li')].map(li => li.innerText.trim())")
    

    现在您可以将它们用于您的逻辑:

    for material in lis:
      if material in lst:
        df.at[code_no, material] = 1
    

    【讨论】:

    • 谢谢你。但是,当我运行它并尝试打印“lis”时,它返回 None 所以我不能在循环中迭代它。您能否解释一下您是如何生成“lis”命令的,以便我尝试修复它?
    • 也许你需要添加一个time.sleep?。本质上,您想在 chrome 中加载该页面并在控制台中摆弄 js,直到它返回您想要的内容。没有看到我能提供的最好帮助的网址。
    • 我会添加一个 time.sleep,现在也试试。至于网址:recyclenow.com/local-recycling - 但是您必须点击 RHS 按钮(“查找最近的回收站”),然后输入邮政编码(尝试“NP11”)并点击搜索。然后单击第一个站点,您将看到我提到的格式。如您所见,Selenium 需要采取很多步骤,所以这也会减慢它的速度!
    • 好的。请把我的建议当作一般性的建议。您需要努力使其适应您的具体情况。
    猜你喜欢
    • 2017-03-10
    • 1970-01-01
    • 1970-01-01
    • 2023-03-12
    • 1970-01-01
    • 2022-06-30
    • 2014-07-27
    • 1970-01-01
    • 2022-07-13
    相关资源
    最近更新 更多