【问题标题】:How to solve StaleElementReferenceException when looping trough elements (Selenium)循环遍历元素时如何解决 StaleElementReferenceException (Selenium)
【发布时间】:2019-12-13 15:01:21
【问题描述】:

我正在尝试通过网络抓取一个网站以获取有关足球比赛的信息。因此我在 Python 中使用 Selenium 库。

我将所有需要的匹配项中的可点击 html 元素存储在一个名为“completed_matches”的列表中。我创建了一个 for 循环,它遍历所有这些可点击的 html 元素。在循环中,我单击当前的 html 元素并打印新的 URL。代码如下所示:

from selenium import webdriver
import selenium
from selenium.webdriver.support.ui import WebDriverWait

driver = webdriver.Chrome(r"C:\Users\Mart\Downloads\chromedriver_win32_2\chromedriver.exe")
url = "https://footystats.org/spain/la-liga/matches"
driver.get(url)
completed_matches = driver.find_elements_by_xpath("""//*[@id="matches-list"]/div[@class='full-matches-table mt2e ' or @class='full-matches-table mt1e ']/div/div[2]/table[@class='matches-table inactive-matches']/tbody/tr[*]/td[3]/a[1]/span""");
print(len(completed_matches))
for match in completed_matches:
        match.click()
        print("Current driver URL: " + driver.current_url)

输出如下:

159
Current driver URL: https://footystats.org/spain/fc-barcelona-vs-real-club-deportivo-mallorca-h2h-stats#632514
---------------------------------------------------------------------------
StaleElementReferenceException            Traceback (most recent call last)
<ipython-input-3-da5851d767a8> in <module>
      4 print(len(completed_matches))
      5 for match in completed_matches:
----> 6         match.click()
      7         print("Current driver URL: " + driver.current_url)

~\Anaconda3\lib\site-packages\selenium\webdriver\remote\webelement.py in click(self)
     78     def click(self):
     79         """Clicks the element."""
---> 80         self._execute(Command.CLICK_ELEMENT)
     81 
     82     def submit(self):

~\Anaconda3\lib\site-packages\selenium\webdriver\remote\webelement.py in _execute(self, command, params)
    631             params = {}
    632         params['id'] = self._id
--> 633         return self._parent.execute(command, params)
    634 
    635     def find_element(self, by=By.ID, value=None):

~\Anaconda3\lib\site-packages\selenium\webdriver\remote\webdriver.py in execute(self, driver_command, params)
    319         response = self.command_executor.execute(driver_command, params)
    320         if response:
--> 321             self.error_handler.check_response(response)
    322             response['value'] = self._unwrap_value(
    323                 response.get('value', None))

~\Anaconda3\lib\site-packages\selenium\webdriver\remote\errorhandler.py in check_response(self, response)
    240                 alert_text = value['alert'].get('text')
    241             raise exception_class(message, screen, stacktrace, alert_text)
--> 242         raise exception_class(message, screen, stacktrace)
    243 
    244     def _value_or_default(self, obj, key, default):

StaleElementReferenceException: Message: stale element reference: element is not attached to the page document
  (Session info: chrome=79.0.3945.79)
  (Driver info: chromedriver=72.0.3626.7 (efcef9a3ecda02b2132af215116a03852d08b9cb),platform=Windows NT 10.0.18362 x86_64)

completed_matches 列表包含 159 个 html 元素,但 for 循环只显示第一个点击的链接,然后抛出 StaleElementReferenceException...

有谁知道如何解决这个问题?

【问题讨论】:

    标签: python selenium for-loop selenium-chromedriver staleelementreferenceexception


    【解决方案1】:

    您要查找的网址在您点击的链接中。您选择单击的父元素。 StaleElementReferenceException 是因为在您单击链接后,页面会更改,呈现在第一个被单击的元素之后的所有元素。

    from selenium import webdriver
    import selenium
    from selenium.webdriver.support.ui import WebDriverWait
    
    driver = webdriver.Chrome(r"C:\Users\Mart\Downloads\chromedriver_win32_2\chromedriver.exe")
    url = "https://footystats.org/spain/la-liga/matches"
    driver.get(url)
    completed_matches = driver.find_elements_by_xpath("""//*[@id="matches-list"]/div[@class='full-matches-table mt2e ' or @class='full-matches-table mt1e ']/div/div[2]/table[@class='matches-table inactive-matches']/tbody/tr[*]/td[3]/a[1]/span""");
    print(len(completed_matches))
    for match in completed_matches:
            #match.click()
            #print("Current driver URL: " + driver.current_url)
            match_parent = match.find_element_by_xpath("..")
            href = match_parent.get_attribute("href")
            print("href: ", href)
    

    【讨论】:

      【解决方案2】:

      单击后,DOM 会刷新,因此会出现 StaleElementReferenceException。所以在 for 循环中再次构建 completed_matches 元素。

      completed_matches = driver.find_elements_by_xpath("""//*[@id="matches-list"]/div[@class='full-matches-table mt2e ' or @class='full-matches-table mt1e ']/div/div[2]/table[@class='matches-table inactive-matches']/tbody/tr[*]/td[3]/a[1]/span""");
      print(len(completed_matches))
      for match in completed_matches:
          completed_matches = driver.find_elements_by_xpath("""//*[@id="matches-list"]/div[@class='full-matches-table mt2e ' or @class='full-matches-table mt1e ']/div/div[2]/table[@class='matches-table inactive-matches']/tbody/tr[*]/td[3]/a[1]/span""");
          match.click()
      

      【讨论】:

        【解决方案3】:

        陈旧的意思是陈旧的、腐烂的、不再新鲜的。陈旧元素是指旧元素或不再可用的元素。假设在 WebDriver 中作为 WebElement 引用的网页上有一个元素。如果 DOM 发生变化,那么 WebElement 就会过时。

        这意味着您正在工作的页面在单击该元素后正在发生变化,因此我的建议是修复它:

        from selenium.webdriver.support.ui import WebDriverWait
        from selenium.webdriver.support import expected_conditions as EC
        while True:
             try:
                 completed_match = WebDriverWait(driver,10).until(EC.element_to_be_clickable((By.XPATH, "//*[@id="matches-list"]/div[@class='full-matches-table mt2e ")))
             except TimeoutException:
                   break
             completed_match.click()
             time.sleep(2)
        

        所以只需遍历元素并每次更新它,在这种情况下它肯定会在页面的 DOM 中

        您可以在此处查看带有完整详细信息代码的旅行顾问的网络爬虫:

        https://github.com/alirezaznz/Tripadvisor-Webscraper

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2018-11-02
          • 1970-01-01
          • 1970-01-01
          • 2021-10-12
          • 1970-01-01
          • 1970-01-01
          • 2010-11-16
          相关资源
          最近更新 更多