【问题标题】:Using R2HTML with rvest/xml2将 R2HTML 与 rvest/xml2 一起使用
【发布时间】:2015-06-23 01:57:19
【问题描述】:

我正在阅读this 关于新包 XML2 的博文。以前,rvest 曾经依赖于XML,它通过将两个包中的函数组合起来(至少)使我的很多工作变得更容易:例如,当我不能使用 XML 包中的htmlParse 时使用 html 读取 HTML 页面(现在他们称为 read_html)。

以this 为例,然后我可以在解析的页面上使用rvest 之类的html_nodes、html_attr 函数。现在,rvest 取决于XML2 这是不可能的(至少在表面上)。

我只是想知道 XML 和 XML2 之间的基本区别是什么。除了在前面提到的post中注明XML包的作者外,包的作者并没有解释XML和XML2之间的区别。

另一个例子:

library(R2HTML) #save page as html and read later
library(XML)
k1<-htmlParse("https://stackoverflow.com/questions/30897852/html-in-rvest-verses-htmlparse-in-xml")
head(getHTMLLinks(k1),5) #This works

[1] "//stackoverflow.com"           "http://chat.stackoverflow.com" "http://blog.stackoverflow.com" "//stackoverflow.com"          
[5] "http://meta.stackoverflow.com"

# But, I want to save HTML file now in my working directory and work later

HTML(k1,"k1") #Later I can work with this
rm(k1)
#read stored html file k1
head(getHTMLLinks("k1"),5)#This works too 

[1] "//stackoverflow.com"           "http://chat.stackoverflow.com" "http://blog.stackoverflow.com" "//stackoverflow.com"          
[5] "http://meta.stackoverflow.com"

#with read_html in rvest package, this is not possible (as I know)
library(rvest)
library(R2HTML)
k2<-read_html("https://stackoverflow.com/questions/30897852/html-in-rvest-verses-htmlparse-in-xml")

#This works
df1<-k2 %>%
html_nodes("a")%>%
html_attr("href")

head(df1,5)
[1] "//stackoverflow.com"           "http://chat.stackoverflow.com" "http://blog.stackoverflow.com" "//stackoverflow.com"          
[5] "http://meta.stackoverflow.com"

# But, I want to save HTML file now in my working directory and work later
HTML(k2,"k2") #Later I can work with this
rm(k2,df1)
#Now extract webpages by reading back k2 html file
#This doesn't work
k2<-read_html("k2") 

df1<-k2 %>%
html_nodes("a")%>%
html_attr("href")

df1
character(0)

更新:

#I have following versions of packages loaded: 
lapply(c("rvest","R2HTML","XML2","XML"),packageVersion)
[[1]]
[1] ‘0.2.0.9000’

[[2]]
[1] ‘2.3.1’

[[3]]
[1] ‘0.1.1’

[[4]]
[1] ‘3.98.1.2’

我使用的是 Windows 8、R 3.2.1 和 RStudio 0.99.441。

【问题讨论】:

  • 这对我来说很好。我刚刚安装了rvest 的最新开发版本。也许你也应该更新你的 (R2HTML_2.3.1,rvest_0.2.0.9000,xml2_0.1.1)
  • @MrFlick:我使用的版本和你的一样。它运行,但正如您在帖子中看到的那样,它提供character(0) 作为输出。
  • 我无法复制。我假设你从 github repo 安装了 rvest。在开发过程中,版本号似乎没有改变。我仍然建议您尝试从 repo 重新安装以获取最新版本。此外,也许发布您使用的操作系统、R 版本(基本上是您的sessionInfo(),以防万一这是边缘情况)。
  • 我已经更新了帖子。是的,你是对的;我已经从 github 安装了rvest。
  • 尝试定义HTML.xml_document&lt;-function(x, ...) HTML(as.character(x),...) 或者更好,为什么要使用R2HTML 而不是仅仅用write_xml() 写出数据?

标签: xml r rvest


【解决方案1】:

R2HTML 包似乎只是 XML 对象上的capture.out,然后将其写回磁盘。这似乎不是将 HTML/XML 数据保存回磁盘的可靠方法。两者可能不同的原因是XML 数据的打印输出不同于xml2 数据。你可以定义一个函数来调用as.character(),而不是依赖capture.output

HTML.xml_document<-function(x, ...) HTML(as.character(x),...)

或者您可能完全跳过R2HTML 并直接使用write_xml 写出xml2 数据。

也许最好的方法是先下载文件,然后再导入。

download.file("http://stackoverflow.com/questions/30897852/html-in-rvest-verses-htmlparse-in-xml", "local.html")
k2 <- read_html("local.html")

【讨论】:

  • 谢谢。 download.file 是个不错的选择(从来没想过)。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-07-11
  • 2020-05-15
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多