【问题标题】:cheerio webcrawler get sequence elementCheerio webcrawler 获取序列元素
【发布时间】:2018-11-23 19:53:14
【问题描述】:

我正在开发一个网络爬虫来读取这样的 html 代码:

<h3>title 1</h3>
<p>content 1</p>
<h3>title 2</h3>
<p>content 2</p>
<h3>title 3</h3>
<p>content 3</p>
<h3>title 4</h3>
<p>content 4</p>
<h3>title 5</h3>
<p>content 5</p>

我想将标题 1 与内容 1 匹配,将标题 2 与内容 2 匹配,然后继续。我没有在cheerio 文档或jquery 中找到获取下一个元素或循环所有DOM 的方法。

在文档中,我只能进入元素(子元素)并返回(父元素)。但我找不到下一个'

' 在找到它上面的 '' 之后。

有什么想法吗?

谢谢!

【问题讨论】:

    标签: node.js web-crawler cheerio


    【解决方案1】:

    这里有几种方法:

    const cheerio = require('cheerio')
    const $ = cheerio.load('<h3>title 1</h3><p>content 1</p><h3>title 2</h3><p>content 2</p><h3>title 3</h3><p>content 3</p><h3>title 4</h3><p>content 4</p><h3>title 5</h3><p>content 5</p>')
    
    $('h3').get().map( h3 => {
      let title = $(h3).text()
      let content = $(h3).next().text()
      // or
      content = $(h3.nextSibling).text()
      console.log(title, content)
    } )
    

    jQuery 可以让你做$(h3).find('+ p'),这很好但是cheerio 不支持它。

    【讨论】:

      猜你喜欢
      • 2018-12-10
      • 2017-09-07
      • 1970-01-01
      • 1970-01-01
      • 2020-05-13
      • 1970-01-01
      • 1970-01-01
      • 2021-06-19
      • 2020-07-06
      相关资源
      最近更新 更多