【问题标题】:How to scrape the yahoo finance search auto suggestion result with selenium python?如何使用 selenium python 抓取雅虎财经搜索自动建议结果?
【发布时间】:2018-10-25 13:52:00
【问题描述】:

我正在尝试使用 selenium python 在 yahoo Finance 上进行自动搜索。当我输入一些单词时,会弹出一个建议,就像在谷歌建议上一样。

https://finance.yahoo.com/

我发现一个带有 xpath 的列表元素应该是 yahoo 提出的建议:

//*[@id="search-assist-input"]/div[2]/ul

建议内容似乎隐藏在这个列表中,但它是不可见的,我的意思是当我点击展开它时,它就消失了。我不知道 Firefox 或 chrome 中是否存在某种“始终展开节点”,但这些元素似乎很难达到。 我试图获取该元素下的所有子元素,它显示找不到元素:

from chrome_driver.chrome import Chrome

driver = Chrome().get_driver()
driver.get('https://finance.yahoo.com/')
driver.find_elements_by_xpath("//div[@id='search-assist-input']/div/input")[0].send_keys('goog')
x = driver.find_elements_by_xpath("//div[@data-reactid='56']/ul[@data-reactid='57']/*")

如何从搜索框中找到这些自动建议?

【问题讨论】:

    标签: python selenium selenium-webdriver xpath webdriverwait


    【解决方案1】:

    请在下面找到针对雅虎财经最新变化的修订版本。

    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support import expected_conditions as EC
    
    options = Options()
    options.add_argument("start-maximized")
    options.add_argument("disable-infobars")
    options.add_argument("--disable-extensions")
    
    options.page_load_strategy = 'eager'
    options.add_argument('--ignore-certificate-errors')
    options.add_argument('--ignore-ssl-errors')
    options.add_argument('log-level=3')
    latest_news = ['Go to Latest News']
    
    chrome_path = "C:\Python\SYS\chromedriver.exe"
    driver = webdriver.Chrome(chrome_options=options, executable_path=chrome_path)
    driver.get('https://finance.yahoo.com/')
    
    WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.XPATH, "//input[@name='yfin-usr-qry']"))).send_keys("goog")
    WebDriverWait(driver, 20).until(EC.text_to_be_present_in_element((By.XPATH,'//*[@id="header-search-form"]/div[2]/div[1]/div/div[1]/h3'),'Symbols'))
    
    yahoo_fin_auto_suggestions = driver.find_elements(By.CLASS_NAME,'modules_list__1zFHY')[0].text.split('\n')
    if yahoo_fin_auto_suggestions == latest_news:
        yahoo_fin_auto_suggestions = driver.find_elements(By.CLASS_NAME,'modules_list__1zFHY')[1].text.split('\n')
    
    
    print(yahoo_fin_auto_suggestions)
    
    driver.quit()
    

    【讨论】:

    • 您的答案可以通过额外的支持信息得到改进。请edit 添加更多详细信息,例如引用或文档,以便其他人可以确认您的答案是正确的。你可以找到更多关于如何写好答案的信息in the help center。
    【解决方案2】:

    由于https://finance.yahoo.com/网站可能的源代码已经更改,我将@DebanjanB的答案调整为三点:

    1. 点击接受cookies/提交同意
    2. 搜索字段的 Xpath(至少对于德国/欧盟而言)
    3. 建议列表的 Xpath
    from selenium import webdriver
    from selenium.webdriver.chrome.options import Options
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support import expected_conditions as EC
    
    options = Options()
    options.add_argument("start-maximized")
    options.add_argument("disable-infobars")
    options.add_argument("--disable-extensions")
    #options.add_argument('headless') #optional for headless driver
    
    driver = webdriver.Chrome(chrome_options=options, executable_path=r'C:\Program Files (x86)\Google\Chrome\Chromedriver\chromedriver.exe')
    driver.get('https://finance.yahoo.com/')
    driver.find_element_by_xpath("//button[@type='submit' and @value='agree']").click() #for cookie consent
    
    WebDriverWait(driver, 5).until(EC.element_to_be_clickable((By.XPATH, "//input[@name='yfin-usr-qry']"))).send_keys("goog")
    yahoo_fin_auto_suggestions = WebDriverWait(driver, 20).until(EC.visibility_of_all_elements_located((By.XPATH, '(//div[@class="_0ea0377c _4343c2a0 _50f34a35"])')))
    for item in yahoo_fin_auto_suggestions:
        print(item.text)
    

    【讨论】:

      【解决方案3】:

      提取关于搜索文本的自动建议,例如GOOG 在https://finance.yahoo.com/ 的搜索框 中,您必须诱导WebDriverWait 以使自动建议可见 和您可以使用以下解决方案:

      • 代码块:

        from selenium import webdriver
        from selenium.webdriver.chrome.options import Options
        from selenium.webdriver.support.ui import WebDriverWait
        from selenium.webdriver.common.by import By
        from selenium.webdriver.support import expected_conditions as EC
        
        options = Options()
        options.add_argument("start-maximized")
        options.add_argument("disable-infobars")
        options.add_argument("--disable-extensions")
        driver = webdriver.Chrome(chrome_options=options, executable_path=r'C:\WebDrivers\ChromeDriver\chromedriver_win32\chromedriver.exe')
        driver.get('https://finance.yahoo.com/')
        WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.XPATH, "//input[@name='p']"))).send_keys("goog")
        yahoo_fin_auto_suggestions = WebDriverWait(driver, 20).until(EC.visibility_of_all_elements_located((By.XPATH, "//input[@name='p']//following::div[1]/ul//li")))
        for item in yahoo_fin_auto_suggestions :
            print(item.text)
        
      • 控制台输出:

        GOOG
        Alphabet Inc.Equity - NASDAQ
        GOOGL
        Alphabet Inc.Equity - NASDAQ
        GOOGL-USD.SW
        AlphabetEquity - Swiss
        GOOGL180518C01080000
        GOOGL May 2018 call 1080.000Option - OPR
        GOOG.MX
        Alphabet Inc.Equity - Mexico
        GOOG180525C01075000
        GOOG May 2018 call 1075.000Option - OPR
        GOOG180518C00720000
        GOOG May 2018 call 720.000Option - OPR
        GOOGL180518C01120000
        GOOGL May 2018 call 1120.000Option - OPR
        GOOGL.MX
        Alphabet Inc.Equity - Mexico
        GOOGL190621C01500000
        GOOGL Jun 2019 call 1500.000Option - OPR
        

      【讨论】:

      • 很好,请问您是如何知道它存储在 li 节点中的?我应该如何使用 firefox/chrome 检查它?
      • @bot1 这些<li> 标签是基于ReactJS 的标签,你可以随时取出page_source 来检查内容。作为一个捷径,我拍了一张快照来识别子标签是<li>标签:)
      • 好的,但是当我点击展开 ul 节点时,它就消失了。你是如何为它拍摄快照的?我的意思是,在你点击展开后,你是不是很快就完成了?
      • @bot1 我不应该一开始就推出捷径 :) 让我们遵循最佳实践,即page_source
      • 是的,我实际上刚刚意识到这正是我在另一个问题中所寻找的。​​span>
      猜你喜欢
      • 2023-03-27
      • 1970-01-01
      • 1970-01-01
      • 2020-09-10
      • 1970-01-01
      • 2020-02-24
      • 2017-01-14
      • 1970-01-01
      • 2020-04-01
      相关资源
      最近更新 更多