【问题标题】:Web Scraping in R | Unable to extract information under a certain node using rvestR 中的网页抓取 |无法使用rvest提取某个节点下的信息
【发布时间】:2019-10-29 10:10:14
【问题描述】:

我正在尝试从网站 (here) 中提取节点 /html/head/script[16] 下的一些信息,但我无法这样做。

nykaa <- "https://www.nykaa.com/biotique-bio-kelp-protein-shampoo-for-falling-hair-intensive-hair-growth-treatment-conf/p/357142?categoryId=1292&productId=357142&ptype=product&skuId=39934"

obj <- read_html(nykaa)

extracted_json <- obj %>% 
  html_nodes(xpath = "/html/head/script[16]") %>% 
  html_text(trim = TRUE)

目前,上述代码的输出为空。但我想有条不紊地提取上述节点下的数据。

【问题讨论】:

  • 看起来您在调用html_node 时使用的xpath 不好。您要提取哪些信息?
  • @ulfelder 我正在尝试从该节点中存在的信息中提取价格、数量等特征。而且我无法提取我指定的 Xpath 之下的 Xpath!
  • 如果我运行html_nodes(obj, xpath = "//html//head//script"),我会得到一组只有 4 个节点。所以没有script[16]。
  • @ulfelder 这很奇怪。如果我在 Inspect Element 中搜索“数量”,然后复制存在数量的标签的 Xpath,我会得到“/html/head/script[16]”
  • 我不是这个特定过程的专家,但我认为问题可能在于页面使用 Javascript 来呈现您所追求的信息。所以你从rvest 得到的是对 Javascript 的调用,而不是结果。运行extracted_json &lt;- obj %&gt;% html_nodes(xpath = "//html//head//script") %&gt;% html_text(),然后看看extracted_json[3],看看我的意思。

标签: r web-scraping rvest


【解决方案1】:

您可以使用正则表达式来获取该脚本标记内的 javascript 对象,然后传递给 jsonlite 并进行解析。您需要扎根一点才能从中获得想要的东西,但这一切都在那里

library(rvest)
library(magrittr)
library(stringr)
library(jsonlite)

p <- read_html('https://www.nykaa.com/biotique-bio-kelp-protein-shampoo-for-falling-hair-intensive-hair-growth-treatment-conf/p/357142?categoryId=1292&productId=357142&ptype=product&skuId=39934') %>% html_text()
all_data <- jsonlite::parse_json(str_match_all(p,'window\\.__PRELOADED_STATE__ = (.*)')[[1]][,2])

【讨论】:

    猜你喜欢
    • 2021-12-12
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-02-15
    • 1970-01-01
    • 2017-10-02
    • 2020-07-18
    • 2019-02-17
    相关资源
    最近更新 更多