【发布时间】:2015-01-27 11:07:30
【问题描述】:
我想从网站 url 上抓取手机的价格:http://www.flipkart.com/apple-iphone-5s/p/itmdv6f75dyxhmt4?pid=MOBDPPZZDX8WSPAT
如果查看代码,价格放在下面的SPAN中
<div class="pricing line">
<div class="prices" itemprop="offers" itemscope="" itemtype="http://schema.org/Offer">
<div>
<span class="selling-price omniture-field" data-omnifield="eVar48" data-eVar48="37500">Rs. 37,500</span> // Fetch this price
</div>
<span class="sticky-message">Selling Price</span>
<meta itemprop="price" content="37,500">
<meta itemprop="priceCurrency" content="INR">
</div>
</div>
到目前为止,我获取此代码的代码是:
<?php
$curl = curl_init('http://www.flipkart.com/apple-iphone-5s/p/itmdv6f75dyxhmt4?pid=MOBDPPZZDX8WSPAT');
curl_setopt($curl, CURLOPT_RETURNTRANSFER, TRUE);
$page = curl_exec($curl);
if(!empty($curl)){ //if any html is actually returned
$pokemon_doc->loadHTML($curl);
libxml_clear_errors(); //remove errors for yucky html
$pokemon_xpath = new DOMXPath($pokemon_doc);
//get all the h2's with an id
$pokemon_row = $pokemon_xpath->query('//h2[@id]');
if($pokemon_row->length > 0){
foreach($pokemon_row as $row){
echo $row->nodeValue . "<br/>";
}
}
}
else
print "Not found";
?>
这显示一个错误:
致命错误:在非对象上调用成员函数 loadHTML() D:\xampp\htdocs\jiteen\php-scrape\phpScrape.php 在第 9 行
我该怎么办,我无法追踪错误
【问题讨论】:
-
我可以建议simplehtmldom.sourceforge.net 相信我,它很棒。并且非常易于使用。
-
您好@kkaosninja,感谢您的帮助和时间。但老实说,我不太能够满足我的要求(可能是因为我没有仔细阅读文档)。你能建议我一个短而快的方法吗?此外,代码很难理解我从那里下载的文件:simple_html_dom.html。
标签: php html xpath web-scraping domdocument