【发布时间】:2020-07-01 00:01:10
【问题描述】:
现在我的这个代码可以点击这个页面的每个产品“https://www.daraz.com.bd/audio/?page=1&spm=a2a0e.home.cate_2.2.49c74591NNpWDU%27”,它把我带到每个项目的产品详细信息页面。谁能告诉我如何循环多个页面,例如 page2、page3、page4?这是我的代码
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
from bs4 import BeautifulSoup
#argument for incognito Chrome
option = Options()
option.add_argument("--incognito")
browser = webdriver.Chrome(options=option)
browser.get("https://www.daraz.com.bd/audio/?page=1&spm=a2a0e.home.cate_2.2.49c74591NNpWDU%27")
# Wait 20 seconds for page to load
timeout = 20
try:
WebDriverWait(browser, timeout).until(EC.visibility_of_element_located((By.XPATH, "//div[@class='c16H9d']")))
except TimeoutException:
print("Timed out waiting for page to load")
browser.quit()
soup = BeautifulSoup(browser.page_source, "html.parser")
product_items = soup.find_all("div", attrs={"data-qa-locator": "product-item"})
for item in product_items:
item_url = f"https:{item.find('a')['href']}"
print(item_url)
browser.get(item_url)
item_soup = BeautifulSoup(browser.page_source, "html.parser")
# Use the item_soup to find details about the item from its url.
container = item_soup.find_all("div",attrs={"id":"container"})
for items in container:
title = items.find("div",{"class":"pdp-product-title"})
print(title)
browser.quit()
现在它只从 page1 获取信息。我希望它也会从其他页面收集信息,例如 page2,page3,page4,page5
【问题讨论】:
-
每个页面都有一个分页对象,可以带你到上一页/下一页等
-
我会怎么做?如果可能,请用代码解释。感谢您的评论
-
@FarhanAhmed Stack Overflow 不能替代指南、教程或文档。
标签: python selenium selenium-webdriver web-scraping beautifulsoup