【问题标题】:R-Advanced Web Scraping-bypassing aspNetHidden using xmlTreeParse()R-Advanced Web Scraping-绕过 aspNetHidden 使用 xmlTreeParse()
【发布时间】:2014-01-26 22:28:07
【问题描述】:

这个问题需要一点时间来介绍,请多多包涵。如果你能到达那里,解决它会很有趣。这种抓取将使用循环复制到该网站上的数千个页面。

我正在尝试抓取网站http://www.digikey.com/product-detail/en/207314-1/A25077-ND/,以获取带有 Digi-Key 零件编号、可用数量等的表格中的数据。包括右侧的价格中断、单价、扩展价格。

使用 R 函数 readHTMLTable() 不起作用,只会返回 NULL 值。这样做的原因(我相信)是因为该网站在 html 代码中使用标签“aspNetHidden”隐藏了它的内容。

出于这个原因,我还发现在使用 htmlTreeParse() 和 xmlTreeParse() 时遇到困难,因为整个部分都没有出现在结果中。

使用scrapeR package中的R函数scrape()

require(scrapeR)

URL<-scrape("http://www.digikey.com/product-detail/en/207314-1/A25077-ND/")

返回完整的 html 代码,包括感兴趣的行:

<th align="right">Digi-Key Part Number</th>
<td id="reportpartnumber">
<meta itemprop="productID" content="sku:A25077-ND">A25077-ND</td>

<th>Price Break</th>
<th>Unit Price</th>
<th>Extended Price
</th>
</tr>
<tr>
<td align="center">1</td>
<td align="right">2.75000</td>
<td align="right">2.75</td>

但是,我无法从这段代码中选择节点,并返回错误:

no applicable method for 'xpathApply' applied to an object of class "list"

我使用不同的函数收到了该错误,例如:

xpathSApply(URL,'//*[@id="pricing"]/tbody/tr[2]')

getNodeSet(URL,"//html[@class='rd-product-details-page']")

我对 xpath 不是最熟悉,但一直在使用网页上的检查元素识别 xpath 并复制 xpath。

非常感谢您对此提供的任何帮助!

【问题讨论】:

  • scrape 来自哪个包?
  • 我猜它来自 scrapeR
  • 我认为该站点实际上会在用户代理字符串中的任何位置检查“wget”,如果存在则返回错误页面。邪恶。

标签: html r xpath web-scraping scrape


【解决方案1】:

你还没有读过scrape的帮助吗?它返回一个列表,您需要获取该列表的一部分(如果 parse=TRUE)等等。

我还认为该网页正在执行一些繁重的浏览器检测。如果我尝试从命令行输入wget 页面,我会得到一个错误页面,scrape 函数会得到一些可用的东西(但对你来说似乎不同),Chrome 会得到所有编码内容的完整垃圾。呸。这对我有用:

> URL<-scrape("http://www.digikey.com/product-detail/en/207314-1/A25077-ND/")
> tables = xpathSApply(URL[[1]],'//table')
> tables[[2]]
<table class="product-details" border="1" cellspacing="1" cellpadding="2">
  <tr class="product-details-top"/>
  <tr class="product-details-bottom">
    <td class="pricing-description" colspan="3" align="right">All prices are in US dollars.</td>
  </tr>
  <tr>
    <th align="right">Digi-Key Part Number</th>
    <td id="reportpartnumber"><meta itemprop="productID" content="sku:A25077-ND"/>A25077-ND</td>
    <td class="catalog-pricing" rowspan="6" align="center" valign="top">
      <table id="pricing" frame="void" rules="all" border="1" cellspacing="0" cellpadding="1">
        <tr>
          <th>Price Break</th>
          <th>Unit Price</th>
          <th>Extended Price&#13;
</th>
        </tr>
        <tr>
          <td align="center">1</td>
          <td align="right">2.75000</td>
          <td align="right">2.75</td>

根据您的用例调整,在这里我获取所有表格并显示第二个表格,其中包含您想要的信息,其中一些在您可以直接获取的 pricing 表格中:

pricing = xpathSApply(URL[[1]],'//table[@id="pricing"]')[[1]]

> pricing
<table id="pricing" frame="void" rules="all" border="1" cellspacing="0" cellpadding="1">
  <tr>
    <th>Price Break</th>
    <th>Unit Price</th>
    <th>Extended Price&#13;
</th>
  </tr>
  <tr>
    <td align="center">1</td>
    <td align="right">2.75000</td>
    <td align="right">2.75</td>
  </tr>

等等。

【讨论】:

  • 很好的答案,感谢您的快速回复。将定价表转换为数据框的最佳方式是什么?
猜你喜欢
  • 1970-01-01
  • 2021-09-03
  • 2020-04-08
  • 2017-08-31
  • 2020-07-21
  • 1970-01-01
  • 1970-01-01
  • 2017-10-31
相关资源
最近更新 更多