【问题标题】:Checking if Xpath Exists检查 Xpath 是否存在
【发布时间】:2018-08-22 09:20:52
【问题描述】:

我目前正在使用 selenium、bs4 和 python 进行抓取,但是在检查 Xpath 是否存在时遇到了问题,这是我的代码:

def hasXpath(xpath):
    try:
        browser.get(quote_page) 
        self.browser.find_element_by_xpath(xpath)
        return True
    except:
        return False                          

# IF PRICELIST EXISTS CONDITION/    

if hasXpath("(//div[@id='product-header h4']//span)[last()-2]") or hasXpath("(//div[@id='product-header']//span)[last()-1]") or hasXpath("(//div[@id='product-header']//span)[last()]"):
    #No. of Items Per Retail(NEED LOGIN)
    somthing = browser.find_element_by_xpath("(//div[@id='product-header']//td)[21]").get_attribute("innerText")
    print(somthing)

    #Retail Price(NEED LOGIN)
    browser.get(quote_page)
    somthing1 = browser.find_element_by_xpath("(//div[@id='product-header']//span)[last()]").get_attribute("innerText")
    print(somthing1)


    if hasXpath("(//div[@id='product-header']//span)[last()-1]"):

       #1No. of Items Per Retail(NEED LOGIN)
       something = browser.find_element_by_xpath("(//div[@id='product-header']//td)[19]").get_attribute("innerText")
       print(something)

       #1Retail Price(NEED LOGIN)
       browser.get(quote_page) 
       something = browser.find_element_by_xpath("(//div[@id='product-header']//span)[last()-1]").get_attribute("innerText")
       print(something1)
else:
  print("It didn't go inside")

如您所见,它有一个简单的函数hasXpath(),我在它下面的 IF 语句中传递条件的 xpath。但是,当我测试它时,一切似乎都在 else 语句中。我还尝试将 True 条件加倍,但没有运气。实施时我做错了什么?

【问题讨论】:

  • 添加一个 minimal reproducible example,其中包含 XPath 失败的标记的缩减副本。没有这个,你的问题是不完整的。
  • checking the Xpath if it exists or not 的概念基本上是模糊的。也许用户编写 xpath 来定位元素。最多您可以检查一个元素(由 xpath 标识)是否陈旧。

标签: python html selenium xpath beautifulsoup


【解决方案1】:

问题

您似乎在使用一个不存在的变量:self。 (至少,您的代码 sn-p 没有任何迹象表明它可能存在或这些可能是实例方法。)

如果您删除包装 try: ... except: ...,我敢打赌您会看到以下错误:

NameError: name 'self' is not defined

解决办法

假设 browser 已定义,只需删除 self.。

奖励 1:更好的异常处理

一般来说,您应该:

  • 尽可能缩小您预计会失败的范围(try: 块中包含的内容)和
  • 尽可能具体地说明您处理的异常情况。

在您的情况下,您似乎只想保护 find_element_by_xpath() 并且只想捕获 Selenium 的 NoSuchElementException:

browser.get(quote_page)
try:
    browser.find_element_by_xpath(xpath)
    return True
except NoSuchElementException:
    return False

这种特殊性将允许使用self 引起的NameError 冒泡到您可能看到的位置,从而省去了在此处发布问题的麻烦。

奖励 2:更高效的页面加载

您调用browser.get(quote_page) 多达六次:在您的脚本正文中两次;一次在hasXpath函数中,被调用了四次。

由于您总是加载相同的页面,因此只需在脚本开始时执行一次。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-03-02
    相关资源
    最近更新 更多