【问题标题】:scraping text from a list of urls using r使用 r 从 url 列表中抓取文本
【发布时间】:2018-04-12 03:02:00
【问题描述】:

我正在尝试使用 r 从 url 列表中删除标题和内容。我能够单独提取每篇文章的标题和内容。但是,我需要遍历这些 url 列表以从每个页面及其内容中获取标题。

这些是 url,它们存储在 csv 文件中: http://well.blogs.nytimes.com/2016/08/29/edible-sunscreens-all-the-rage-but-no-proof-they-work/?smid=fb-nytwell&smtyp=cur

http://www.nytimes.com/2016/08/30/well/live/how-12-epipens-saved-my-life.html?smid=fb-nytwell&smtyp=cur

http://www.nytimes.com/2016/08/29/opinion/why-we-never-die.html?smid=fb-nytwell&smtyp=cur

http://www.nytimes.com/2016/08/31/health/how-to-ride-downhill-on-a-bicycle.html?smid=fb-nytwell&smtyp=cur

http://www.cbssports.com/college-football/news/one-sweet-gesture-by-fsus-travis-rudolph-makes-mom-of-an-autistic-boy-cry/

http://www.nytimes.com/2016/08/31/well/family/what-kids-wish-their-teachers-knew.html?smid=fb-nytwell&smtyp=cur

这是我用来单独提取每篇文章的代码(请注意,内容的每一段都被认为是一个节点,当我提取这些节点时,每个节点都会出现在一个新的原始文件中,而我需要它们只在第一个原始的)。

install.packages('xml2')    
library(xml2)    
library(rvest)

url <- "http://well.blogs.nytimes.com/2016/08/29/edible-sunscreens-all-the-rage-but-no-proof-they-work/?smid=fb-nytwell&smtyp=cur"

article <- read_html(url)    
title <- article %>% html_node(".entry-title") %>% html_text()    
content <- article %>% html_nodes(".story-body-text") %>% html_text()    
article_table <- data.frame(title, content)

article_table

【问题讨论】:

    标签: r


    【解决方案1】:

    您需要折叠输出以进入文章的单行

    content <-
        article %>% html_nodes(".story-body-text") %>% html_text() %>% paste(., collapse = "")
    

    对于多个网址,已经回答here

    根据您的情况对其进行了调整。请注意,.entry-title 标签不适用于所有网址。你需要使用title

    library(rvest)
    library(purrr)
    article <- listofurls %>% map(read_html)
    title <-
        article %>% map_chr(. %>% html_node("title") %>% html_text())
    content <-
        article %>% map_chr(. %>% html_nodes(".story-body-text") %>% html_text() %>% paste(., collapse = ""))
    article_table <- data.frame("Title" = title, "Content" = content)
    dim(article_table)
    

    【讨论】:

    • 太棒了。这很好用。关于遍历这些 url 以提取内容的任何想法?
    • 我认为已经回答了。如果我的解决方案有效,请接受答案。谢谢
    • 感谢 user5249203 的回答。我仍在寻找问题后半部分的答案(如何遍历 url 列表)。我会等待其他人的帮助。
    • 太棒了!这很有帮助。但是对于某些 url,内容不会出现。
    • 例如第3个url和第5个url的内容不显示。
    猜你喜欢
    • 2021-05-13
    • 2013-01-08
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-07-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多