【发布时间】:2021-12-21 06:05:08
【问题描述】:
使用 Selenium 和 BeautifulSoap,我正在尝试抓取网页。一般来说,这很好用。请在下面找到代码。
在此页面上列出了一些类别。深度为4级。在每个级别我有 20 个项目/链接。
我的问题是:在循环中打开和处理这些链接的最有效方法是什么?
import sys
sys.path.insert(0,'/usr/lib/chromium-browser/chromedriver')
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from bs4 import BeautifulSoup
import time
options = webdriver.ChromeOptions()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
wd = webdriver.Chrome('chromedriver',options=options)
wd.get("url")
source = wd.page_source
soup = BeautifulSoup(source, "html.parser")
items = soup.select('ul[data-card-id="tree-list0972"]')
for item in items:
ul = item.find('ul')
for li in ul:
print(li.a.get('href') + ',' + li.a.text)
cats = webdriver.Chrome('chromedriver',options=options)
options = webdriver.ChromeOptions()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
# Here i do need to open the link from the url list (3 levels deep)
cats.get(h + domain + li.a.get('href'))
WebDriverWait(webdriver, timeout=3)
cats.close
wd.close
【问题讨论】:
-
为什么首先需要打开单独的浏览器实例? (除了代码中的其他问题)
-
嗯,打开网址?我想你想说页面源存储在变量中,通过循环变量项我不需要打开浏览器的新实例,对吧?
-
我是说用同一个浏览器实例打开网址。
-
你想怎么做?对于每个链接,您想在同一个浏览器选项卡中打开还是希望在新选项卡/浏览器实例中打开?另外,您总共有多少链接?一个链接已经加载,你要检索什么吗?
-
网址是什么?很可能数据可以通过 API 获得?
标签: python selenium-webdriver beautifulsoup selenium-chromedriver