从 BeautifulSoup 页面检索所有信息答案

【问题标题】：Retrieving all information from page BeautifulSoup从 BeautifulSoup 页面检索所有信息
【发布时间】：2017-04-08 06:05:14
【问题描述】：

我正在尝试在 OldNavy 网页上抓取产品的网址。然而，它只是给出了产品列表的一部分而不是整个列表（例如，当 URL 超过 8 个时，只给出 8 个）。我希望有人能帮忙找出问题所在。

from bs4 import BeautifulSoup
from selenium import webdriver
import html5lib
import platform
import urllib
import urllib2
import json


link = http://oldnavy.gap.com/browse/category.do?cid=1035712&sop=true
base_url = "http://www.oldnavy.com"

driver = webdriver.PhantomJS()
driver.get(link)
html = driver.page_source
soup = BeautifulSoup(html, "html5lib")
bigDiv = soup.findAll("div", class_="sp_sm spacing_small")
for div in bigDiv:
  links = div.findAll("a")
  for i in links:
    j = j + 1
    productUrl = base_url + i["href"]
    print productUrl

【问题讨论】：

此代码不起作用 - 您有没有 "" 的 url 和 j 的错误。在提出问题之前检查代码。

标签： python selenium-webdriver web-scraping beautifulsoup web-crawler

【解决方案1】：

此页面使用JavaScript 加载元素，但它仅在您向下滚动页面时加载。

它被称为"lazy loading"。

你也必须滚动页面。

from selenium import webdriver
from bs4 import BeautifulSoup
import time

link = "http://oldnavy.gap.com/browse/category.do?cid=1035712&sop=true"
base_url = "http://www.oldnavy.com"

driver = webdriver.PhantomJS()
driver.get(link)

# ---

# scrolling

lastHeight = driver.execute_script("return document.body.scrollHeight")
#print(lastHeight)

pause = 0.5
while True:
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(pause)
    newHeight = driver.execute_script("return document.body.scrollHeight")
    if newHeight == lastHeight:
        break
    lastHeight = newHeight
    #print(lastHeight)

# ---

html = driver.page_source
soup = BeautifulSoup(html, "html5lib")

#driver.find_element_by_class_name

divs = soup.find_all("div", class_="sp_sm spacing_small")
for div in divs:
    links = div.find_all("a")
    for link in links:
    print base_url + link["href"]

想法：https://stackoverflow.com/a/28928684/1832058

【讨论】：