【问题标题】:web scraping with R, clicking on links用 R 抓取网页,点击链接
【发布时间】:2018-08-01 17:53:09
【问题描述】:

我是初学者,我想从页面中抓取所有带有所选关键字的文章。我只能抓取显示在单个页面上的文章标题、文章描述的一部分及其链接。我不仅想抓取搜索结果,还想抓取每个显示链接的内容。

网站:http://search.time.com/?site=time&q=bitcoin

require(rvest)
url<- "http://search.time.com/?site=time&q=bitcoin"
webpage <- read_html(url)

title_data_html <- html_nodes(webpage,'.content-title a')

title_data <- html_text(title_data_html)

description_data_html <- html_nodes(webpage,'.content-snippet')
description_data <- html_text(description_data_html)

links = html_attr(title_data_html, name = "href")

【问题讨论】:

    标签: r web-scraping rvest


    【解决方案1】:

    您所追求的功能是来自rvest 包的follow_link()。这是关于此主题的另一篇 SO 帖子:

    Scraping linked HTML webpages by looping the rvest::follow_link() function

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2019-07-12
      • 2016-03-21
      • 2020-08-08
      • 1970-01-01
      • 1970-01-01
      • 2012-03-28
      • 2021-05-03
      相关资源
      最近更新 更多