【问题标题】:Parsing large XML file in R is very slow在 R 中解析大型 XML 文件非常慢
【发布时间】:2015-08-26 11:08:53
【问题描述】:

我需要从 R 中的一个大型 xml 文件中提取数据。文件大小为 60 MB。我使用以下R代码从网上下载数据:

library(XML)
library(httr)

url = "http://hydro1.sci.gsfc.nasa.gov/daac-bin/his/1.0/NLDAS_NOAH_002.cgi"
SOAPAction = "http://www.cuahsi.org/his/1.0/ws/GetSites"
envelope = "<?xml version=\"1.0\" encoding=\"utf-8\"?>\n<soap:Envelope xmlns:xsi=\"http://www.w3.org/2001/XMLSchema-instance\" xmlns:xsd=\"http://www.w3.org/2001/XMLSchema\" xmlns:soap=\"http://schemas.xmlsoap.org/soap/envelope/\">\n<soap:Body>\n<GetSites xmlns=\"http://www.cuahsi.org/his/1.0/ws/\">\n<site></site><authToken></authToken>\n</GetSites>\n</soap:Body>\n</soap:Envelope>"

response = POST(url, body = envelope,
             add_headers("Content-Type" = "text/xml", "SOAPAction" = SOAPAction))
status.code = http_status(response)$category

收到服务器的响应后,我使用以下代码将数据解析为 data.frame:

# Parse the XML into a tree
WaterML = content(response, as="text")
SOAPdoc = xmlRoot(xmlTreeParse(WaterML, getDTD=FALSE, useInternalNodes = TRUE))
doc = SOAPdoc[[1]][[1]][[1]]

# Allocate a new empty data frame with same name of rows as the number of sites
N = xmlSize(doc) - 1
df = data.frame(SiteName=rep("",N),
             SiteID=rep(NA, N),
             SiteCode=rep("",N),
             Latitude=rep(NA,N),
             Longitude=rep(NA,N),
             stringsAsFactors=FALSE)

# Populate the data frame with the values
# This loop is VERY SLOW it takes around 10 MINUTES!
start.time = Sys.time()

for(i in 1:N){  
  siteInfo = doc[[i+1]][[1]]
  siteList = xmlToList(siteInfo)
  siteName = siteList$siteName
  sCode = siteList$siteCode
  siteCode = sCode$text
  siteID = ifelse(is.null(sCode$.attrs["siteID"]), siteCode,   sCode$.attrs["siteID"])
  latitude = as.numeric(siteList$geoLocation$geogLocation$latitude)
  longitude = as.numeric(siteList$geoLocation$geogLocation$longitude) 
}

end.time = Sys.time()
time.taken = end.time - start.time
time.taken

我用来将 XML 解析为 data.frame 的 for 循环非常慢。大约需要 10 分钟才能完成。有什么方法可以让循环更快?

【问题讨论】:

  • 这是一个非常大的 XML 数据集,因此使用特定于 XML 的库进行解析需要相当长的时间也就不足为奇了。如果数据非常结构化,那么您可以轻松编写自己的循环结构以及一些正则表达式来解析数据。但似乎这是一个一次性问题,所以 10 分钟似乎是一个不错的折衷方案,而解决方案可能需要超过 10 分钟才能解决?
  • 对我来说这不是一次性问题,因为在线 XML 数据集每天都在更新。所以我需要尽可能快地进行解析。
  • 这么多次调用xmlToList的瓶颈是什么(N的典型值是多少?)?您能否将整个 xmldoc 转换为一个列表并使用它? 60Mb 并不是真的“大”(它应该适合 RAM),所以我希望它是可能的,而且可能会更快。
  • 另外,你的代码实际上并没有在循环中改变df,所以它没有填充数据框!将其编写为一个函数,并以 N 作为参数,这样您就可以在更少的行上对其进行测试,从而不必等待 20 分钟来查看您的代码是否有效。
  • 您无需进行xmlToList 转换即可提取元素。尝试按名称访问节点,例如:doc[[123]][[1]][["geoLocation"]][["geogLocation"]][["latitude"]][["text"]] 获取纬度。或者如果您确信格式是恒定的,则按数字(例如:doc[[123]][[1]][[3]][[1]][[1]][["text"]]。此外,在整个数据框列的末尾进行数字转换(df$latitude = as.numeric(df$latitude))。

标签: xml r performance xml-parsing dataframe


【解决方案1】:

通过使用 xpath 表达式来提取您想要的信息,我能够获得更好的性能。在我的笔记本电脑上,对xpathSApply 的每次调用都需要大约 20 秒,因此所有命令都在 2 分钟内完成。

# you need to specify the namespace information
ns <- c(soap="http://schemas.xmlsoap.org/soap/envelope/",
        xsd="http://www.w3.org/2001/XMLSchema",
        xsi="http://www.w3.org/2001/XMLSchema-instance",
        sr="http://www.cuahsi.org/waterML/1.0/",
        gsr="http://www.cuahsi.org/his/1.0/ws/")

Data <- list(
  siteName = xpathSApply(SOAPdoc, "//sr:siteName", xmlValue, namespaces=ns),
  siteCode = xpathSApply(SOAPdoc, "//sr:siteCode", xmlValue, namespaces=ns),
  siteID = xpathSApply(SOAPdoc, "//sr:siteCode", xmlGetAttr, name="siteID", namespaces=ns),
  latitude = xpathSApply(SOAPdoc, "//sr:latitude", xmlValue, namespaces=ns),
  longitude = xpathSApply(SOAPdoc, "//sr:longitude", xmlValue, namespaces=ns))
DataFrame <- as.data.frame(Data, stringsAsFactors=FALSE)
DataFrame$latitude <- as.numeric(DataFrame$latitude)
DataFrame$longitude <- as.numeric(DataFrame$longitude)

【讨论】:

  • 非常好的解决方案,在我的笔记本电脑上花了大约 2 分钟才能完成。并感谢您提供如何使用 xpathSApply 的示例。我怀疑“xpath”函数会更快,但我不知道如何正确指定命名空间信息。您的回答是一个很好的例子,如何在 R 中使用 xpath 来解析具有大量命名空间的 XML 文档。
猜你喜欢
  • 2018-10-13
  • 2013-02-28
  • 2016-11-26
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2011-05-09
  • 2013-03-24
相关资源
最近更新 更多