【问题标题】:Python BeautifulSoup Returning Empty ListPython BeautifulSoup 返回空列表
【发布时间】:2017-02-27 00:21:59
【问题描述】:

我正在尝试创建一个 Python 脚本,以使用 BeautifulSoup 从 tcgplayer.com 提取游戏王卡的价格。当您在此网站上搜索卡片时,它会返回一个搜索结果页面,其中包含来自不同卖家的多个价格。我的目标是拉动所有这些价格。在下面的示例中,我打开了一张名为“A”细胞繁殖设备的卡片的搜索​​结果:

import urllib2
from bs4 import BeautifulSoup
html = urllib2.open('http://shop.tcgplayer.com/productcatalog/product/show?newSearch=false&ProductType=All&IsProductNameExact=false&ProductName=%22A%22%20Cell%20Breeding%20Device')
soup = BeautifulSoup(html, 'lxml')
soup.find_all('span', {'class': 'scActualPrice largetext pricegreen'})

几天前,正确运行soup.find_all 行为我提供了所需的信息。但是,现在运行它会给我一个空数组 []。我已经广泛搜索了关于 BeautifulSoup 返回一个空数组的信息,但我不确定它们中的任何一个是否适用于我,因为它在几天前工作得很好。有人可以帮我指出正确的方向吗?提前谢谢!

【问题讨论】:

    标签: python web-scraping beautifulsoup


    【解决方案1】:

    您应该使用selenium 使用真实浏览器进行报废:

    from selenium import webdriver
    
    driver = webdriver.Chrome('/path/to/chromedriver')
    driver.get('http://shop.tcgplayer.com/productcatalog/product/show?newSearch=false&ProductType=All&IsProductNameExact=false&ProductName=%22A%22%20Cell%20Breeding%20Device')
    prices = driver.find_elements_by_css_selector('.scActualPrice')
    for element in prices:
        print(element.text)
    driver.quit()
    

    【讨论】:

      【解决方案2】:

      本网站使用名为 Incapsula 的服务。网站开发人员配置 Incapsula 以防止机器人访问其内容。

      我建议您联系他们的管理员并请求访问权限或向他们索取 API。

      【讨论】:

      • 使用 selenium 对我有用,但你认为它会在几天后停止工作吗?
      • 使用 selenium,你实际上是在打开浏览器并进行所有操作,所以现在应该没问题。但将来可能会有机会。
      • 而且使用 selenium 不可靠
      猜你喜欢
      • 2019-12-12
      • 2019-03-25
      • 1970-01-01
      • 2021-10-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-12-22
      • 1970-01-01
      相关资源
      最近更新 更多