【问题标题】:How to loop multiple page on selenium python BeautifulSoup如何在 selenium python BeautifulSoup 上循环多个页面
【发布时间】:2020-07-01 00:01:10
【问题描述】:

现在我的这个代码可以点击这个页面的每个产品“https://www.daraz.com.bd/audio/?page=1&spm=a2a0e.home.cate_2.2.49c74591NNpWDU%27”,它把我带到每个项目的产品详细信息页面。谁能告诉我如何循环多个页面,例如 page2、page3、page4?这是我的代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
from bs4 import BeautifulSoup

#argument for incognito Chrome
option = Options()
option.add_argument("--incognito")


browser = webdriver.Chrome(options=option)

browser.get("https://www.daraz.com.bd/audio/?page=1&spm=a2a0e.home.cate_2.2.49c74591NNpWDU%27")


# Wait 20 seconds for page to load
timeout = 20
try:
    WebDriverWait(browser, timeout).until(EC.visibility_of_element_located((By.XPATH, "//div[@class='c16H9d']")))
except TimeoutException:
    print("Timed out waiting for page to load")
    browser.quit()


soup = BeautifulSoup(browser.page_source, "html.parser")

product_items = soup.find_all("div", attrs={"data-qa-locator": "product-item"})
for item in product_items:
    item_url = f"https:{item.find('a')['href']}"
    print(item_url)

    browser.get(item_url)

    item_soup = BeautifulSoup(browser.page_source, "html.parser")
    # Use the item_soup to find details about the item from its url.
    container = item_soup.find_all("div",attrs={"id":"container"})
    for items in container:
        title = items.find("div",{"class":"pdp-product-title"})
        print(title)



browser.quit()

现在它只从 page1 获取信息。我希望它也会从其他页面收集信息,例如 page2,page3,page4,page5

【问题讨论】:

  • 每个页面都有一个分页对象,可以带你到上一页/下一页等
  • 我会怎么做?如果可能,请用代码解释。感谢您的评论
  • @FarhanAhmed Stack Overflow 不能替代指南、教程或文档。

标签: python selenium selenium-webdriver web-scraping beautifulsoup


【解决方案1】:

当您单击 URL 中的第 2 页时,您可以在网站中看到,唯一更改的是页码,因此可以通过循环轻松完成。 要获得更好的代码,您可以创建一个名为 url 的变量并为每个页面更改它:

for page_num in range(1, 10): # change the range as you want to
    url = "https://www.daraz.com.bd/audio/?page={}&spm=a2a0e.home.cate_2.2.49c74591NNpWDU%27".format(page_num)

并将其余代码放入此循环中(最后一行除外)

【讨论】:

  • 感谢 Kurosh Ghanizadeh。有效。请告诉我如何在 csv 文件中获取输出?我需要编写什么代码才能获得输出?
  • 我没有过多地使用 CSV,所以我建议只是阅读一些文档。这不是一件复杂的事情,您可以轻松使用它
猜你喜欢
  • 2019-05-26
  • 1970-01-01
  • 1970-01-01
  • 2019-01-31
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-03-31
相关资源
最近更新 更多