【问题标题】:Selenium TimeoutException: Message:硒超时异常:消息:
【发布时间】:2020-07-19 04:10:12
【问题描述】:

我有以下代码从堆栈溢出中抓取问题。当我将class names 用作css selectors 时,例如:question-hyperlink 是类名,我将其转换为css 选择器为.question-hyperlink。它可以正常工作。

但是当我将tagNames 用作CSS selectors 如.A 或.DIV 时,它返回给我这个超时错误:

Traceback (most recent call last):
  File "C:\Users\intel\Desktop\D_scraper2.pyw", line 37, in to_do
    element = WebDriverWait(driver, 5).until(
  File "C:\Users\intel\AppData\Local\Programs\Python\Python38\lib\site-packages\selenium\webdriver\support\wait.py", line 80, in until
    raise TimeoutException(message, screen, stacktrace)
selenium.common.exceptions.TimeoutException: Message: 

我的代码:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import pandas as pd
import time
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import csv

def to_do():
# vars...
    csv_file_location = r"C:\Users\intel\Desktop\data_file.csv"

    user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_3) AppleWebKit/537.36 (KHTML, like Gecko) ' \
                 'Chrome/80.0.3987.132 Safari/537.36'

    driver_exe = 'chromedriver'
    options = Options()
    options.add_argument("--headless")
    options.add_argument(f'user-agent={user_agent}')
    options.add_argument("--disable-web-security")
    options.add_argument("--allow-running-insecure-content")
    options.add_argument("--allow-cross-origin-auth-prompt")

    url = "https://stackoverflow.com/questions"

    driver = webdriver.Chrome(executable_path=r"C:\Users\intel\Downloads\setups\chromedriver.exe", options=options)
    driver.get(url)

    one_ = ".A"

    two_ = ".DIV"

    three_ = ".A"

    try:
        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, one_))
        )
        elements_1 = driver.find_elements_by_css_selector(one_)

        web_content_list = []
        for ele in elements_1:
            web_content_dict = {}
            web_content_dict["Title"] = ele.text
            web_content_list.append(web_content_dict)

        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, two_))
        )
        elements_2 = driver.find_elements_by_css_selector(two_)

        for ele2 in elements_2:
            web_content_dict = {}
            web_content_dict["Title2"] = ele2.text
            web_content_list.append(web_content_dict)

        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, three_))
        )

        elements_3 = driver.find_elements_by_css_selector(three_)

        for ele3 in elements_3:
            web_content_dict = {}
            web_content_dict["Title3"] = ele3.text
            web_content_list.append(web_content_dict)

        df = pd.DataFrame(web_content_list)
        new_df = pd.DataFrame({'Column 1': df['Title'].dropna(),
                  'Column 2': df['Title2'].dropna(),
                  'Column 3': df['Title3'].dropna()})
        new_df.to_csv(csv_file_location,
                  index=False, mode='a', encoding='utf-8')

        try:
            f = open(csv_file_location)
            print("Done !!!\n"*3)

        except IOError:
            print("File not accessible")

        finally:
            f.close()
        driver.quit()

    except:
        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, one_))
        )
        elements_1 = driver.find_elements_by_css_selector(one_)

        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, two_))
        )

        elements_2 = driver.find_elements_by_css_selector(two_)

        element = WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, three_))
        )

        elements_3 = driver.find_elements_by_css_selector(three_)

        df = pd.DataFrame({
            "Title1" : [ele for ele.text in elements_1],
            "Title2" : [ele2 for ele2.text in elements_2],
            "Title3" : [ele3 for ele3.text in elements_3],
        })
        df.to_csv(csv_file_location,
                  index=False, mode='w', encoding='utf-8')

        try:
            f = open(csv_file_location)
            print("Done !!!\n"*3)
            # Do something with the file
        except IOError:
            print("File not accessible")

        finally:
            f.close()
        driver.quit()

    finally:
        print("start")

if __name__ == "__main__":
    to_do()

在 class names 和 css selectors 的情况下,我在 csv 文件中使用了三列,在 tagNames 的情况下,我使用三列到 csv。也许这些标签中有更多的列......

任何帮助将不胜感激......

【问题讨论】:

    标签: python python-3.x selenium web-scraping css-selectors


    【解决方案1】:

    当您想使用 css 选择器搜索标签时,不要使用它前面的句点。如果你想使用'a'锚标签,你可以这样做:

    one_ = "a"
    
    element = WebDriverWait(driver, 5).until(
                EC.presence_of_element_located((By.CSS_SELECTOR, one_))
            )
    

    将点放在前面用于类名。

    【讨论】:

    • 感谢@RKelley,它运行良好。但是 csv 文件之间有太多空格。它还刮掉了所有不需要的东西。如何限制它?例如,如果我只想获取 h1 标签中的问题,而不是 h1 标签中的 stackoverflow 网站名称
    • 我在回答您在其他地方发布的相同问题后才看到这个。为了限制问题,您应该将 one_ 更改为 '#questions h3 a'。这将做的是搜索“问题”的 id,然后查找任何带有子锚 ('a') 标记的子 h3 元素。
    猜你喜欢
    • 1970-01-01
    • 2021-03-19
    • 1970-01-01
    • 1970-01-01
    • 2020-08-30
    • 2021-08-23
    • 1970-01-01
    • 2022-11-19
    • 1970-01-01
    相关资源
    最近更新 更多