【问题标题】:How to get rid of None while looping through webelements?如何在遍历 webelements 时摆脱 None ?
【发布时间】:2021-02-13 22:59:34
【问题描述】:

我正在尝试返回没有任何 None 的 web 元素列表。不知道为什么,但似乎pass 不起作用。有任何想法吗?请。

顺便说一句。可以使用 Pandas 修复它,但我想坚持使用纯 Python/Selenium 并了解问题所在。

def get_(article):
    try:
        article.find_element_by_xpath(".//a[div[@class='accessible_elem']]")
    except NoSuchElementException:
        pass
    else:
        title = article.find_element_by_xpath(".//a[div[@class='accessible_elem']]").get_attribute('aria-label')
        pubdate = article.find_element_by_xpath(".//abbr").get_attribute('data-utime')
        url = article.find_element_by_xpath(".//a[div[@class='accessible_elem']]").get_attribute('href')
        return(title, pubdate, url)

output = []
for article in articles:
    content = get_(article)
    output.append(content)

【问题讨论】:

  • if title is None: pass,否则...应该可以工作。请注意,一篇文章可能存在,但具有 None 作为您要求的属性的值。
  • 关键元素在 try: 语句中。其余的应该在 True 时执行。当然有些文章没有 class='accessible_elem'。

标签: python selenium exception web-scraping


【解决方案1】:

问题:您的NoSuchElementException 有时不会被捕获,因为没有抛出 NoSuchElementException。一个示例是,如果存在具有类 accessible_elem 的元素,但没有您读取的正确属性。此外,当您因异常而通过时,该函数返回 None。

修复:视情况而定,但您可能想先检查内容是否为无,然后在附加之前检查标题、发布日期或网址中的任何一个是否为无。将您的 for 循环更改为:

for article in articles:
  content = get_(article)
  if content and all([x is not None for x in content]):
    output.append(content)

您可以将检查缩短为:

if content and all(content):

如果你知道你永远不会得到任何元组值的值 0(一个假值)。

【讨论】:

  • 谢谢。稍作修改即可工作:如果内容不是无:
  • 已修改以反映这一点。您可能想从您的 catch 块中显式返回 None
  • pass 和 return 都产生 None 原因不明。
【解决方案2】:

您无意中使用了try-except-else 结构。

建议使用list comprehensions:

def analyze_article(article):
    try:
        article.find_element_by_xpath(".//a[div[@class='accessible_elem']]")
    except NoSuchElementException:
        return None
    
    # we know that the element must exist if we get to this point
    # if it did not, we returned None already and left the function body
    title = article.find_element_by_xpath(".//a[div[@class='accessible_elem']]").get_attribute('aria-label')
    pubdate = article.find_element_by_xpath(".//abbr").get_attribute('data-utime')
    url = article.find_element_by_xpath(".//a[div[@class='accessible_elem']]").get_attribute('href')
    return (title, pubdate, url)


# Get all items, be them None or not
items = [analyze_article(art) for art in articles]

# Filter out all None values
items = [item for item in items if item is not None]

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-05-23
    • 2014-06-18
    • 1970-01-01
    • 2020-09-22
    • 1970-01-01
    相关资源
    最近更新 更多