【问题标题】:Why cant i iterate through a list in selenium?为什么我不能遍历硒列表?
【发布时间】:2021-08-20 14:24:57
【问题描述】:

我正在尝试抓取维基百科页面的标题作为练习,我希望能够区分带有“h2”和“h3”标签的标题。

所以我写了这段代码:

from selenium import webdriver                  
from selenium.webdriver.common.keys import Keys                                             #For being able to input key presses
import time                                                                                 #Useful for if your browser is faster than your code
PATH = r"C:\Users\Alireza\Desktop\chromedriver\chromedriver.exe"                            #Location of the chromedriver
driver = webdriver.Chrome(PATH)

driver.get("https://de.wikipedia.org/wiki/Alpha-Beta-Suche")                                #Open website in Chrome
print(driver.title)                                                                         #Print title of the website to console

h1Header = driver.find_element_by_tag_name("h1")                                            #Find the first heading in the article

h2HeaderTexts = driver.find_elements_by_tag_name("h2")                                          #List of all other major headers in the article

h3HeaderTexts = driver.find_elements_by_tag_name("h3")                                          #List of all minor headers in the article


for items in h2HeaderTexts:
    scor = items.find_element_by_class_name("mw-headline")



driver.quit()

但是,这不起作用并且程序不会终止。 有人有这个解决方案吗?

这里的问题在于for循环!显然我不能从 h2HeaderTexts 中的元素中按类名(或其他任何内容)刮取任何元素,尽管这应该是可能的。

【问题讨论】:

  • 我建议通过添加一些print() 语句来调试您的代码。 h2HeaderTexts的内容是什么?它是一个空列表吗?如果是这样,那么问题在于h2HeaderTexts = driver.find_elements_by_tag_name("h2") 而不是您假设的循环。下一步是配置driver 以打开浏览器窗口,这样您就可以看到它实际抓取的内容。在您尝试获取 <h2> 元素之前,页面可能尚未完全加载。或者该页面可能没有任何 <h2> 元素。或者完全有其他原因。
  • 鉴于您作为答案发布的错误消息,程序确实终止了......它以异常消息终止。问题是,正如 Cruisepandey 指出的那样,第一个 h2(存储在 h2HeaderTexts 中)没有包含类“mw-headline”的后代元素,因此它抛出 NoSuchElementException。使用 CSS 选择器 h2 .mw-headline,而不是两次点击页面,一次用于 H2,另一次用于具有“mw-headline”类的后代。一口气,您将获得您正在寻找的所有元素。
  • 您需要编辑您的问题并添加您作为答案发布的异常消息并澄清问题,因为它确实以异常消息终止。

标签: python selenium selenium-webdriver web web-scraping


【解决方案1】:

您可以在 xpath 中过滤:

PATH = r"C:\Users\Alireza\Desktop\chromedriver\chromedriver.exe"                            #Location of the chromedriver
driver = webdriver.Chrome(PATH)
driver.maximize_window()
driver.implicitly_wait(30)
driver.get("https://de.wikipedia.org/wiki/Alpha-Beta-Suche")                                #Open website in Chrome
print(driver.title)      

for item in driver.find_elements(By.XPATH, "//h2/span[@class='mw-headline']"):
    print(item.text)

这应该给你,h2 标题与 mw-headline 类元素。

输出:

Informelle Beschreibung
Der Algorithmus
Implementierung
Optimierungen
Vergleich von Minimax und AlphaBeta
Geschichte
Literatur
Weblinks
Fußnoten

Process finished with exit code 0

更新 1:

您的循环仍在运行且程序未终止的原因是,如果您查看页面 HTML 源代码,并且第一个 h2 标记,h2 标记没有 child span 和 @987654330 @,所以 selenium 试图定位 HTML DOM 中不存在的元素。你也使用find_elements,如果找到它会返回一个 web 元素列表,如果没有返回一个空列表,这也是你看不到异常的原因。

【讨论】:

  • 嘿,这似乎是一个不错的解决方法。谢谢!虽然我仍然想知道为什么我的版本不起作用...... :(
  • @JuicySamurai :在我的回答中查看上面更新的 1 代码,应该澄清为什么您的版本不起作用。
  • 当 Selenium 试图定位一个不存在的元素时,它会引发异常。这不能解释为什么程序永远不会停止。
【解决方案2】:

您应该等到页面上出现元素后再访问它们。
此外,还有几个带有h1 标签名称的元素。
要搜索元素内的元素,您应该使用以点开头的 xpath。否则,这将搜索整个页面上的第一个匹配项。
该页面上的第一个 h2 元素内部没有类名为 mw-headline 的元素。所以,你也应该处理这个问题。

from selenium import webdriver                  
from selenium.webdriver.common.keys import Keys   
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC                                          #For being able to input key presses
import time                                                                                 #Useful for if your browser is faster than your code
PATH = r"C:\Users\Alireza\Desktop\chromedriver\chromedriver.exe"                            #Location of the chromedriver
driver = webdriver.Chrome(PATH)
wait = WebDriverWait(driver, 20)


driver.get("https://de.wikipedia.org/wiki/Alpha-Beta-Suche")                                #Open website in Chrome
print(driver.title)                                                                         #Print title of the website to console

wait.until(EC.visibility_of_element_located((By.XPATH, "//h1")))

h1Headers = driver.find_elements_by_tag_name("h1")                                            #Find the first heading in the article

h2HeaderTexts = driver.find_elements_by_tag_name("h2")                                          #List of all other major headers in the article

h3HeaderTexts = driver.find_elements_by_tag_name("h3")                                          #List of all minor headers in the article


for items in h2HeaderTexts:
    scor = items.find_elements_by_xpath(".//span[@class='mw-headline']")
    if scor:
        #do what you need with scor[0] element



driver.quit()

【讨论】:

  • 感谢您的帮助,但这也没有帮助我。问题在于“items.find_element_by_class_name(“mw-headline”)”。不知何故,这不起作用。我将发布我的错误代码作为答案。
【解决方案3】:

您的版本未完成执行,因为如果 selenium 找不到元素,它将放弃该进程。

开发人员不喜欢使用 try/catch,但我个人还没有找到更好的解决方法。如果你这样做:

for items in h2HeaderTexts:
    try:
        scor = items.find_element_by_class_name('mw-headline').text
        print(scor)
    except: 
        print('nothing found')

你会注意到它会执行到最后并且你有一个结果。

【讨论】:

  • 如果找不到元素,Selenium 不会“放弃进程”。它抛出一个异常。
猜你喜欢
  • 1970-01-01
  • 2015-12-02
  • 1970-01-01
  • 2017-10-16
  • 1970-01-01
  • 1970-01-01
  • 2012-09-27
  • 1970-01-01
  • 2023-01-05
相关资源
最近更新 更多