【问题标题】:How to get next node after parent of current via symfony crawler?如何通过 symfony 爬虫获取当前父节点之后的下一个节点?
【发布时间】:2017-05-11 07:18:01
【问题描述】:

用于解析的示例 HTML 5:

<div id="orderDetails">
    <div> ... any number of blocks with unnecessary stuff ... </div>
    <div>Label for important info</div>
    <table> ... some other block type ... </table>
    <div>Some very important info here</div>
    <div> ... any number of blocks with unnecessary stuff ... </div>
</div>

我的 PHP 代码如下所示:

$crawler = new \Symfony\Component\DomCrawler\Crawler($html);
$label = $crawler->filter('#orderDetails div:contains("Label for important info")');
$info = $label->parent()->next('div');
assert('Some very important info here' === $info->text(), 'Important info must be grabbed from HTML');

但不幸的是爬虫没有方法parent和next。但是..它有parents,它给了我所有个父节点==我不能区分的所有div。

所以在这种情况下我有两个问题:

  1. 如何获取当前节点的父节点?不是所有节点,而是“实际”一个!
  2. 如何使用next/prev的一些类似物水平遍历dom?

谢谢。

【问题讨论】:

  • 我注意到 the API mentions parents() “返回当前选择的父母”。不知道对你有没有帮助?
  • 是的,我看到了,但我只需要“当前父节点,而不是父节点s”和“只有一个下一个节点,而不是 全部他们”。但是,在阅读了源代码(旨在重新发明轮子)之后,我发现它只是方法的名称非常令人困惑(请参阅下面的答案)。

标签: php html dom web-crawler symfony


【解决方案1】:

故事

在深入研究源代码后,我发现nextAll() 方法返回的不是“全部”而是“一个”节点 ($node = $this-&gt;getNode(0);)。

这意味着如果我需要“当前之后的两个节点”,那么我必须写$node-&gt;nextAll()-&gt;nextAll()-&gt;nextAll()。

WTF?!这是超级奇怪的命名约定(0_0)。

答案

  1. 如何获取当前节点的父节点?不是所有节点,而是“实际”一个!
// This is only one parent node
$parent = $node->parents();
  1. 如何使用类似 next/prev 的方式水平遍历 dom?
// This is only one node – next after current
$next = $node->nextAll();
// This is only one node – previous before current
$prev = $node->nextAll();
// This is only one node – next after two from current
$nextAfterTwo = $node->nextAll()->nextAll()->nextAll();

具体代码解决方案

因此,由于确实存在所需的实现,问题的功能解决方案如下所示:

/**
 * Returns sibling node that is after current and filtered with selector
 *
 * @param Crawler $start    Node from which start traverse
 * @param string  $selector CSS/XPath selector like in `Crawler::filter($selector)`
 *
 * @return Crawler Found node wrapped with Crawler
 *
 * @throws \InvalidArgumentException When node not found
 */
function getNextFiltered(Crawler $start, string $selector) : Crawler
{
    $count = $start->parents()->count();
    $next  = $start->nextAll();
    while ($count --> 0) {
        $filtered = $next->filter($selector);
        if ($filtered->count()) return $filtered;
        $next = $next->nextAll();
    }

    throw new \InvalidArgumentException('No node found');
}

在我的例子中:

$crawler = new Crawler($html);
$label   = $crawler->filter('#orderDetails div:contains("Label for important info")');
$info    = getNextFiltered($label, 'div');
assert('Some very important info here' === $info->text(), 'Important info must be grabbed from HTML');

【讨论】:

  • 如果(例如)您需要在当前节点之后的第三个节点,您可以使用$node-&gt;nextAll()-&gt;eq(2)。这意味着$node-&gt;nextAll() 将产生与$node-&gt;nextAll()-&gt;eq(0) 相同的结果,后者在当前节点之后获取下一个节点。
猜你喜欢
  • 1970-01-01
  • 2011-10-24
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-09-17
  • 2012-02-14
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多