【发布时间】:2014-01-26 22:28:07
【问题描述】:
这个问题需要一点时间来介绍,请多多包涵。如果你能到达那里,解决它会很有趣。这种抓取将使用循环复制到该网站上的数千个页面。
我正在尝试抓取网站http://www.digikey.com/product-detail/en/207314-1/A25077-ND/,以获取带有 Digi-Key 零件编号、可用数量等的表格中的数据。包括右侧的价格中断、单价、扩展价格。
使用 R 函数 readHTMLTable() 不起作用,只会返回 NULL 值。这样做的原因(我相信)是因为该网站在 html 代码中使用标签“aspNetHidden”隐藏了它的内容。
出于这个原因,我还发现在使用 htmlTreeParse() 和 xmlTreeParse() 时遇到困难,因为整个部分都没有出现在结果中。
使用scrapeR package中的R函数scrape()
require(scrapeR)
URL<-scrape("http://www.digikey.com/product-detail/en/207314-1/A25077-ND/")
返回完整的 html 代码,包括感兴趣的行:
<th align="right">Digi-Key Part Number</th>
<td id="reportpartnumber">
<meta itemprop="productID" content="sku:A25077-ND">A25077-ND</td>
<th>Price Break</th>
<th>Unit Price</th>
<th>Extended Price
</th>
</tr>
<tr>
<td align="center">1</td>
<td align="right">2.75000</td>
<td align="right">2.75</td>
但是,我无法从这段代码中选择节点,并返回错误:
no applicable method for 'xpathApply' applied to an object of class "list"
我使用不同的函数收到了该错误,例如:
xpathSApply(URL,'//*[@id="pricing"]/tbody/tr[2]')
getNodeSet(URL,"//html[@class='rd-product-details-page']")
我对 xpath 不是最熟悉,但一直在使用网页上的检查元素识别 xpath 并复制 xpath。
非常感谢您对此提供的任何帮助!
【问题讨论】:
-
scrape来自哪个包? -
我猜它来自 scrapeR
-
我认为该站点实际上会在用户代理字符串中的任何位置检查“wget”,如果存在则返回错误页面。邪恶。
标签: html r xpath web-scraping scrape