【发布时间】:2021-08-20 14:24:57
【问题描述】:
我正在尝试抓取维基百科页面的标题作为练习,我希望能够区分带有“h2”和“h3”标签的标题。
所以我写了这段代码:
from selenium import webdriver
from selenium.webdriver.common.keys import Keys #For being able to input key presses
import time #Useful for if your browser is faster than your code
PATH = r"C:\Users\Alireza\Desktop\chromedriver\chromedriver.exe" #Location of the chromedriver
driver = webdriver.Chrome(PATH)
driver.get("https://de.wikipedia.org/wiki/Alpha-Beta-Suche") #Open website in Chrome
print(driver.title) #Print title of the website to console
h1Header = driver.find_element_by_tag_name("h1") #Find the first heading in the article
h2HeaderTexts = driver.find_elements_by_tag_name("h2") #List of all other major headers in the article
h3HeaderTexts = driver.find_elements_by_tag_name("h3") #List of all minor headers in the article
for items in h2HeaderTexts:
scor = items.find_element_by_class_name("mw-headline")
driver.quit()
但是,这不起作用并且程序不会终止。 有人有这个解决方案吗?
这里的问题在于for循环!显然我不能从 h2HeaderTexts 中的元素中按类名(或其他任何内容)刮取任何元素,尽管这应该是可能的。
【问题讨论】:
-
我建议通过添加一些
print()语句来调试您的代码。h2HeaderTexts的内容是什么?它是一个空列表吗?如果是这样,那么问题在于h2HeaderTexts = driver.find_elements_by_tag_name("h2")而不是您假设的循环。下一步是配置driver以打开浏览器窗口,这样您就可以看到它实际抓取的内容。在您尝试获取<h2>元素之前,页面可能尚未完全加载。或者该页面可能没有任何<h2>元素。或者完全有其他原因。 -
鉴于您作为答案发布的错误消息,程序确实终止了......它以异常消息终止。问题是,正如 Cruisepandey 指出的那样,第一个 h2(存储在
h2HeaderTexts中)没有包含类“mw-headline”的后代元素,因此它抛出NoSuchElementException。使用 CSS 选择器h2 .mw-headline,而不是两次点击页面,一次用于 H2,另一次用于具有“mw-headline”类的后代。一口气,您将获得您正在寻找的所有元素。 -
您需要编辑您的问题并添加您作为答案发布的异常消息并澄清问题,因为它确实以异常消息终止。
标签: python selenium selenium-webdriver web web-scraping