【问题标题】:Scraping dynamically generated data with scrapy (and selenium?) [duplicate]用scrapy(和selenium?)抓取动态生成的数据[重复]
【发布时间】:2015-12-14 07:31:05
【问题描述】:

我正在努力让 scrapy(有或没有 selenium)从网页中提取动态生成的内容。该网站列出了不同大学的表现,并允许您选择该大学提供的每个学习领域。例如,从下面代码中列出的页面中,我希望能够提取大学名称(“邦德大学”)和“总体体验质量”的值(91.3%)。

但是,当我使用“查看源代码”、curl 或 scrapy 时,不会显示实际值。例如。我希望看到 Uni 名称的地方,它显示:

<h1 class="inline-block instiution-name" data-bind="text: Description"></h1>

但如果我使用 firebug 或 chrome 来检查元素,它会显示

<h1 class="inline-block instiution-name" data-bind="text: Description">Bond University</h1>

经过进一步检查,在 firebug 的“Net”选项卡上,我可以看到有一个 AJAX (?) 调用正在返回相关信息,但我无法在 scrapy 甚至 curl 中模仿这一点(是的,我确实搜索并花了很长时间尝试,恐怕)。

请求标头

POST /Websilk/DataServices/SurveyData.asmx/FetchInstitutionStudyAreaData HTTP/1.1
Host: www.qilt.edu.au
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.8; rv:39.0) Gecko/20100101 Firefox/39.0
Accept: application/json, text/javascript, */*; q=0.01
Accept-Language: en-US,en;q=0.5
Accept-Encoding: gzip, deflate
Content-Type: application/json; charset=utf-8
X-Requested-With: XMLHttpRequest
Referer: http://www.qilt.edu.au/institutions/institution/bond-university/business-management
Content-Length: 36
Cookie: _ga=GA1.3.69062787.1442441726; ASP.NET_SessionId=lueff4ysg3yvd2csv5ixsc1f; _gat=1
Connection: keep-alive
Pragma: no-cache
Cache-Control: no-cache

随请求传递的POST参数

{"InstitutionId":20,"StudyAreaId":0}

作为第二种选择,我尝试将 Selenium 与 scrapy 一起使用,因为我认为它可能会像浏览器一样“看到”真实值,但无济于事。到目前为止,我的主要尝试如下:

import scrapy
import time  #used for the sleep() function

from selenium import webdriver

class QiltSpider(scrapy.Spider):
    name = "qilt"

    allowed_domains = ["qilt.edu.au"]
    start_urls = [
        "http://www.qilt.edu.au/institutions/institution/rmit-university/architecture-building/"
    ]

    def __init__(self):
        self.driver = webdriver.Firefox()
        self.driver.get('http://www.qilt.edu.au/institutions/institution/rmit-university/architecture-building/')
        time.sleep(5) # tried pausing, in case problem was delayed loading - didn't work

    def parse(self, response):
        # parse the response to find the uni name and show in console (using xpath code from firebug). This find the relevant section, but it shows as empty
        title = response.xpath('//*[@id="bd"]/div[2]/div/div/div[1]/div/div[2]/h1').extract()
        print title
        # dumping the whole response to a file so I can check whether dynamic values were captured
        with open("extract.html", 'wb') as f:
            f.write(response.body)
            self.driver.close()

谁能告诉我如何做到这一点?

非常感谢!

编辑:感谢您迄今为止的建议,但是关于如何专门模拟 AJAX 调用 with 的机构 ID 和 StudyAreaID 参数的任何想法?我的测试代码如下,但它似乎仍然遇到错误页面。

import scrapy
from scrapy.http import FormRequest

class HeaderTestSpider(scrapy.Spider):
    name = "headerTest"

    allowed_domains = ["qilt.edu.au"]
    start_urls = [
        "http://www.qilt.edu.au/institutions/institution/rmit-university/architecture-building/"
    ]

    def parse(self, response):
        return [FormRequest(url="http://www.qilt.edu.au/Websilk/DataServices/SurveyData.asmx/FetchInstitutionData",
                            method='POST',  
                            formdata={'InstitutionId':'20', 'StudyAreaId': '0'},
                            callback=self.parser2)]

【问题讨论】:

  • 您可以使用 Requests 库并模仿 AJAX 调用。
  • 这里不需要requests,因为正在使用 Scrapy。
  • 为什么不直接使用 Selenium 并在浏览器中呈现数据后从页面上刮取数据?
  • 解决方案在这里:stackoverflow.com/a/24373576/2368836 您可能需要在 driver.get 之后添加一个隐式等待
  • 感谢您的回复。我之前确实阅读了另一个论坛,并尝试复制 Badarau Petru 使用的方法。关于您对中间件的建议,除了在 settings.py 中启用它之外,我不确定如何模仿 AJAX 调用,特别是使用 InstitutionID 和 StudyAreaID 的参数。不幸的是,我没有经常玩 python,所以它可能超出了我的范围。我实际上创建了一个 elance 工作,看看我是否可以从他们提出的代码中学习。

标签: python selenium web-scraping scrapy


【解决方案1】:

QILT 页面使用 AJAX 从服务器检索数据。此 AJAX 请求使用 javascript 代码发送,该代码使用偶数 document.ready(jQuery)/window.onload(Javascript) 触发(如果您不熟悉 javascript,则在网页完成加载后立即触发此方法浏览器窗口)。由于您使用软件来刺激页面请求,因此根本不会触发此事件。

对于您尝试模拟的 AJAX 请求,请求正文的类型为 Application/JSON。 请在请求中添加以下标头。 内容类型:application/json

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-12-25
    • 1970-01-01
    • 2019-10-27
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多