【问题标题】:How to extract the value of the datetime attribute in a time tag using xpath or css selector in python?如何在 python 中使用 xpath 或 css 选择器提取时间标签中 datetime 属性的值?
【发布时间】:2019-06-03 12:21:05
【问题描述】:

我需要从 HTML 文档中标签的 datetime 属性中提取评论的日期。

我一直在尝试使用 xpath 和 css 选择器的不同变体来实现这一点,但它们返回空字符串。

HTML 标记如下所示:

<time class="review-date--tooltip-target" datetime="2013-10-09T13:47:14.000Z" title= "Wednesday, 9 October 2013, 13:47:14">9 Oct 2013</time>

还有,这是我的 xpath 和 css 选择器:

xpath('//time[@class="review-date--tooltip-target"]')

css('time.review-date--tooltip-target')

两个结果将帮助我:

1- extract the value of the `datetime` attribute

2- extract the text `9 Oct 2013` within the time tag

【问题讨论】:

    标签: python-3.x xpath web-scraping scrapy


    【解决方案1】:

    对于 Scrapy,您需要:

    datetime = response.xpath('//time[@class="review-date--tooltip-target"]/@datetime').extract_first()
    time = response.xpath('//time[@class="review-date--tooltip-target"]/text()').extract_first()
    

    【讨论】:

      【解决方案2】:

      要获取日期时间属性,xpath 表达式

      //time[@class="review-date--tooltip-target"]/@datetime
      

      输出

      2013-10-09T13:47:14.000Z
      

      要获取时间标签内的日期文本,xpath 表达式

      //time[@class="review-date--tooltip-target"]/text()
      

      输出

      9 Oct 2013
      

      【讨论】:

        【解决方案3】:

        尝试以下代码,这应该会返回您的预期值。

        print(driver.find_element_by_xpath("//time[@class='review-date--tooltip-target']").text)
        print(driver.find_element_by_xpath("//time[@class='review-date--tooltip-target']").get_attribute("datetime"))
        

        输出:

        9 Oct 2013
        2013-10-09T13:47:14.000Z
        

        或者你可以诱导WebdriverWait

        from selenium import webdriver
        from selenium.webdriver.common.by import By
        from selenium.webdriver.support.ui import WebDriverWait
        from selenium.webdriver.support import expected_conditions as EC
        element=WebDriverWait(driver,20).until(EC.element_to_be_clickable((By.XPATH,"//time[@class='review-date--tooltip-target']")))
        print(element.text)
        print(element.get_attribute("innerHTML"))
        print(element.get_attribute("datetime"))
        

        或者您可以尝试使用 python Beautifulsoup 进行scraping。

        from selenium import webdriver
        from bs4 import BeautifulSoup
        driver=webdriver.Chrome()
        driver.get("URL")
        html=driver.page_source
        soup=BeautifulSoup(html,'html.parser')
        print(soup.find('time').text)
        print(soup.find('time')['datetime'])
        

        通过使用scrapy选择器try that.get()将返回第一个匹配如果有多个匹配尝试使用getall()


        Datetimeval = response.css('time::attr(datetime)').get()
        Textval = response.css('time::text').get()
        

        【讨论】:

        • 感谢您如此迅速地回复!我可能应该提到我正在使用scrapy 的Selector 对象。它仍然没有返回日期。出于兴趣,您为什么认为相同的 xpath 模式在使用选择器而不是 selenium 时不起作用?
        • @LalehH :好的,我以为您想使用 Python selenium 提取值。
        猜你喜欢
        • 1970-01-01
        • 2018-08-08
        • 2023-03-21
        • 2010-12-30
        • 2020-01-02
        • 1970-01-01
        • 2023-01-29
        • 2021-10-23
        • 1970-01-01
        相关资源
        最近更新 更多