【问题标题】:simple_html_dom - read html page, two arrayssimple_html_dom - 读取 html 页面,两个数组
【发布时间】:2016-08-12 04:11:06
【问题描述】:

这是我的全部代码

// include the scrapper 
include('simple_html_dom.php');

// connect the page for scrapping
$html = file_get_html('http://www.niagarafallsreview.ca/news/local');

// make empty arrays
$headlines = array();
$links = array();

// look for 'h' headings on page
foreach($html->find('h1') as $header) {
    $headlines[] = $header->plaintext;
}

// look for 'a' links that start with 'http://www.niagarafallsreview.ca/2016/04/'
foreach($html->find('a[href^="http://www.niagarafallsreview.ca/2016/04/"]')  as $link) {
    $links[] = $link->href;
}

// trim the headlines because one on top and bottom were not needed
$output = array_slice($headlines, 1, -1); 

// for each header output a nice list of the headers 
foreach ($output as $headers){
    echo "< a href='#'>$headers</a>" . "<br />";
}

// make sure the links are unique and no doubles are found
$result = array_unique($links);

// for each link output it in a nice list
foreach ($result as $linkk){
    echo "<a href='$linkk'>$linkk</a>" . "<br />";
}   

此代码将在一个不错的列表中生成标题,并且还将生成一个不错的链接列表。

我的问题是我需要将它们组合起来,我希望 $header 是 href 的文本,并且 href 中的链接是 $linkk

像这样..

< a href ='$linkk'>$headers</a>

我不知道该怎么做,因为我有两个 foreach 语句。我尝试将它们结合起来,但没有成功。

任何帮助将不胜感激。

谢谢。

【问题讨论】:

  • 也许你显示源 html?为什么你认为数组有相同的长度并且一一对应?
  • 来源是连接页面的第 5 行以下。如果我自己回显数组,它们就会被正确索引
  • 你能解决这个问题吗?我们的回答有帮助吗?

标签: php html html-parsing simple-html-dom


【解决方案1】:

这是您要查找的 foreach:

foreach($output as $i=>$headers) {
  $linkk = $result[$i];

  echo "< a href='$linkk'>$headers</a>" . "<br />";
}

这假设数组具有相同的长度和正确的顺序。

【讨论】:

    【解决方案2】:

    试试这个:

    // include the scrapper 
    include('simple_html_dom.php');
    
    // connect the page for scrapping
    $html = file_get_html('http://www.niagarafallsreview.ca/news/local');
    
    // make empty arrays
    $headlines = array();
    $links = array();
    
    // look for 'h' headings on page
    foreach($html->find('h1') as $header) {
        $headlines[] = $header->plaintext;
    }
    
    // look for 'a' links that start with 'http://www.niagarafallsreview.ca/2016/04/'
    foreach($html->find('a[href^="http://www.niagarafallsreview.ca/2016/04/"]')  as $link) {
        $links[] = $link->href;
    }
    
    // trim the headlines because one on top and bottom were not needed
    $output = array_slice($headlines, 1, -1); 
    
    // make sure the links are unique and no doubles are found
    $result = array_unique($links);
    
    // for each link output it in a nice list
    foreach ($result as $i=>$linkk) {
        $headline = isset($output[$i]) ? $output[$i] : '(empty)';
        echo "<a href='$linkk'>$headline</a>" . "<br />";
    }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-08-08
      • 2019-08-05
      • 2020-02-09
      • 2015-12-10
      • 2013-06-18
      • 1970-01-01
      相关资源
      最近更新 更多