【问题标题】:Getting text within <a > tag inside <p> tag在 <p> 标签内的 <a> 标签内获取文本
【发布时间】:2021-10-21 06:29:50
【问题描述】:

您好,我一直在尝试将 div - p 标签中的所有文本部分添加到 hr 标签,所以有人给了这个 xpath

//div[@class="entry"]/*[not(preceding-sibling::hr | self::hr)]/text()

可以正常工作,但这会忽略 p 标签中 <.a> 标记中的文本部分 有什么想法可以获取该文本吗?

<div class="entry">
   <p> some text</p>
   <p> some text2</p>
   <p> some text3</p>
   <p> some text4
       <a href='somelink'> this text here i want to get through xpath</a>
       some text5
   </p>
   <hr>(up to this hr tag)
   <p> some text5</p>
   <hr>
   <p> some text6</p>
</div>

【问题讨论】:

    标签: html web-scraping xpath scrapy


    【解决方案1】:

    一种方法可能是//div[@class="entry"]/*[not(preceding-sibling::hr | self::hr)]//text(),尽管我可能更喜欢简单地选择元素//div[@class="entry"]/*[not(preceding-sibling::hr | self::hr)] 并使用字符串值。

    【讨论】:

    • 谢谢你,我为什么没有尝试添加 // 但它也适用于我我想知道你描述的第二种方法是什么?
    • 我在做的是 response.xpath('xpath here//text()').get() or .extract()
    【解决方案2】:

    您可以简单地基于 xpath 拉取数据。

    //div[@class="entry"]/p[0]
    //div[@class="entry"]/p[1]
    //div[@class="entry"]/p[2]
    //div[@class="entry"]/p[3]
    //div[@class="entry"]/p[4]
    //div[@class="entry"]/p[5]
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2021-09-07
      • 1970-01-01
      • 1970-01-01
      • 2021-10-09
      • 2012-02-15
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多