【发布时间】:2018-04-12 03:02:00
【问题描述】:
我正在尝试使用 r 从 url 列表中删除标题和内容。我能够单独提取每篇文章的标题和内容。但是,我需要遍历这些 url 列表以从每个页面及其内容中获取标题。
这些是 url,它们存储在 csv 文件中: http://well.blogs.nytimes.com/2016/08/29/edible-sunscreens-all-the-rage-but-no-proof-they-work/?smid=fb-nytwell&smtyp=cur
http://www.nytimes.com/2016/08/29/opinion/why-we-never-die.html?smid=fb-nytwell&smtyp=cur
这是我用来单独提取每篇文章的代码(请注意,内容的每一段都被认为是一个节点,当我提取这些节点时,每个节点都会出现在一个新的原始文件中,而我需要它们只在第一个原始的)。
install.packages('xml2')
library(xml2)
library(rvest)
url <- "http://well.blogs.nytimes.com/2016/08/29/edible-sunscreens-all-the-rage-but-no-proof-they-work/?smid=fb-nytwell&smtyp=cur"
article <- read_html(url)
title <- article %>% html_node(".entry-title") %>% html_text()
content <- article %>% html_nodes(".story-body-text") %>% html_text()
article_table <- data.frame(title, content)
article_table
【问题讨论】:
标签: r