【问题标题】:PHP - Regex/Function for Node Traversing DOM to get specific tagPHP - 用于节点遍历 DOM 以获取特定标签的正则表达式/函数
【发布时间】:2022-06-23 02:22:37
【问题描述】:

我正在使用 Goutte 通过 PHP 抓取 URL。

我想在这个标签之后保存一个列表<ul>...</ul>: <p><strong>Maladies fréquentes :</strong></p>

DOM 看起来像这样的结构:

<p>....</p>
<p>....</p>
<p>....</p>
<p>....</p>
...
<h2>...</h2>
...
<ul>...</ul>
...
<p><strong>Maladies fréquentes :</strong></p>
<ul>
<li>Text I need</li>
<li>Text I need</li>
</ul>
...
<p></p>
<p></p>
...

实际上,我使用:first-of-type 保存到我的数据库中

$crawler->filter('.desc ul:first-of-type li')->each(function ($node) use (&$out) {

   $li = array();

   if ($node->count() > 0) {
        $li[] = str_replace('"', "'", trim($node->filter('li')->text()));
   }

   // Insert into DV

}

当内容包含2个或3个&lt;ul&gt;...&lt;/ul&gt;时总是保存错误的li,因为所有的ul都被选中了。

如何在&lt;p&gt;&lt;strong&gt;Maladies fréquentes :&lt;/strong&gt;&lt;/p&gt; 之后只选择&lt;ul&gt;?

谢谢!

【问题讨论】:

    标签: php web-crawler nodes goutte


    【解决方案1】:

    对Goutte不太了解,但相信你可以将爬虫对象加载到DomDocument中,然后用xpath解析。比如:

    $doc = new DOMDocument();    
    $doc->loadHTML($crawler);
    #or possibly: $doc->loadHTML((string)$crawler);
    $xpath = new DOMXPath($doc);
    $targets = $xpath->query('//p[strong]/following-sibling::ul[1]//li');
    foreach ($targets as $source) {
        echo($source->nodeValue."\r\n");
    };
    

    输出应该是

    Text I need
    Text I need
    

    【讨论】:

      猜你喜欢
      • 2017-09-21
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多