【问题标题】:Beautifulsoup parse thousand pagesBeautifulsoup 解析千页
【发布时间】:2017-09-08 13:16:43
【问题描述】:

我有一个脚本可以解析包含数千个 url 的列表。 但我的问题是,这份清单需要很长时间才能完成。

URL 请求大约需要 4 秒才能加载页面并可以被解析。有什么方法可以快速解析大量的 URL?

我的代码如下所示:

from bs4 import BeautifulSoup   
import requests                 

#read url-list
with open('urls.txt') as f:
    content = f.readlines()
# remove whitespace characters
content = [line.strip('\n') for line in content]

#LOOP through urllist and get information
for i in range(5):
    try:
        for url in content:

            #get information
            link = requests.get(url)
            data = link.text
            soup = BeautifulSoup(data, "html5lib")

            #just example scraping
            name = soup.find_all('h1', {'class': 'name'})

编辑: 在这个例子中如何处理带有钩子的异步请求?我尝试了本网站上提到的以下内容Asynchronous Requests with Python requests:

from bs4 import BeautifulSoup   
import grequests

def parser(response):
    for url in urls:

        #get information
        link = requests.get(response)
        data = link.text
        soup = BeautifulSoup(data, "html5lib")

        #just example scraping
        name = soup.find_all('h1', {'class': 'name'})

#read urls.txt and store in list variable
with open('urls.txt') as f:
    urls= f.readlines()
# you may also want to remove whitespace characters 
urls = [line.strip('\n') for line in urls]

# A list to hold our things to do via async
async_list = []

for u in urls:
    # The "hooks = {..." part is where you define what you want to do
    # 
    # Note the lack of parentheses following do_something, this is
    # because the response will be used as the first argument automatically
    rs = grequests.get(u, hooks = {'response' : parser})

    # Add the task to our list of things to do via async
    async_list.append(rs)

# Do our list of things to do via async
grequests.map(async_list, size=5)

这对我不起作用。我什至没有在控制台中收到任何错误,它只是运行了很长时间直到它停止。

【问题讨论】:

  • 文档描述了如何使用事件挂钩。只需将您的内容处理代码放入一个函数并将其绑定到response 事件。这对requests.get() 和async.get() 的工作方式完全相同。
  • 我编辑了代码,这样怎么用?对不起,我发现初学者的文档有点差。但我也注意到,这在 Python3 中不可用。我用 Python3 编写了整个脚本
  • 不完全,请参阅此处了解工作代码:stackoverflow.com/questions/9110593/…。根据同一线程中第二个答案的 cmets,您可以在 Python3 上使用 async io for HTTP。这篇博文似乎详细解决了这个问题:pawelmhm.github.io/asyncio/python/aiohttp/2016/04/22/…
  • 很抱歉没有把它写成答案,但是写一个有用的答案比我今天要花更多的时间。我相信您可以将博客文章中的代码转换为解决方案 - 一旦您弄清楚了,请发布您自己的答案。我会留下来投票。

标签: python-3.x web-scraping beautifulsoup python-requests grequests


【解决方案1】:

如果有人对这个问题感到好奇 - 我决定从零开始我的项目并使用 scrapy 而不是 beautifulsoup。

Scrapy 是一个完整的网络抓取框架,它具有一次处理 1000 个请求的内置功能,您可以限制您的请求,以便从目标站点“更友好”地抓取。

我希望这可能对某人有所帮助。对我来说,这是这个项目的更好选择。

【讨论】:

    猜你喜欢
    • 2012-10-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-09-23
    • 2020-12-26
    • 2023-03-10
    相关资源
    最近更新 更多