【问题标题】:Scraping contents of web that has specific url抓取具有特定 url 的网页内容
【发布时间】:2012-07-22 02:24:18
【问题描述】:

我想获取具有特定文档链接的标题和网址。所以,从下面的代码中,我想获取信息:标题和http://linkWeb.com具有特定的url下载.pdf http://link.pdf

这是html页面:

<div class="title-download">
<div id="01divTitle" class="title">
    <h3>
        <a id="01Title" onmousedown="" href="http://linkWeb.com">Titles</a>
        <span id="01LbCitation" class="citation">(<a id="01Citation" href="http://citation.com">Citations</a>)</span></h3>
</div>
<div id="01downloadDiv" class="download">
    <a id="01_downloadIcon" title="http://link.pdf" onmousedown="" target=""><img id="ctl01_icon" class="small-icon";" /></a>
</div>

这是代码,但它返回空白结果:

<?php
include 'simple_html_dom.php';
set_time_limit(0);
$url  ='http://example.com';
$html = file_get_html($url) or die ('invalid url');

foreach($html->find('span[class=citation]') as $link){
    foreach($link->parent()->parent()->find('.download a') as $link2){  //I confused with the code in this line
       if(strtolower(substr($link2->title, strrpos($link2->title, '.'))) === '.pdf') {
           $link = $link->prev_sibling();
           echo $link->plaintext.'<br>';
           echo $link->href.'<br>';
       echo $link2->title.'<br>'; 
       }
    }
}
?>

【问题讨论】:

  • 等等,http://link.pdf?这是如何运作的?或者这只是一个虚拟 URL,而不是发布实际的站点名称?
  • @Matchu 哦,对不起。 html页面中的错字。我编辑了它 = title"http://link.pdf"

标签: php simple-html-dom


【解决方案1】:

鉴于 $link 是引用跨度,$link-&gt;parent()-&gt;parent() 返回 ID 为 01divTitle 的 div。而且,由于div 是您正在寻找的.download 元素的兄弟,而不是父元素,$link-&gt;parent()-&gt;parent()-&gt;find('.download a') 不会返回任何结果。

也许$link-&gt;parent()-&gt;parent()-&gt;parent()-&gt;find('.download a') 会更好。可能还有其他问题,但这绝对是其中之一。

【讨论】:

  • 感谢您的解释。它运作良好。非常感谢你:)
猜你喜欢
  • 2013-09-04
  • 2014-05-08
  • 2010-10-09
  • 1970-01-01
  • 2019-01-13
  • 1970-01-01
相关资源
最近更新 更多