【问题标题】:Regex that catches all relative and absolute links from a html string从 html 字符串中捕获所有相对和绝对链接的正则表达式
【发布时间】:2017-05-22 07:27:37
【问题描述】:

我需要捕获来自多个网站的所有链接。为此,我收集了整个 html 文件。我需要一个将所有这些都放在一个数组中的正则表达式。

我不想收集任何图像文件或其他代码文件。只是页面本身的 html。


我希望它收集所有这样的链接:

/https://www.hello.com
/https://www.hello.com/index.php
/https://www.hello.com/world
/https://www.hello.com/world.php
/https://www.hello.com/world.html
/https://hello.com
/https://hello.com/world
/http://www.hello.com
/http://www.hello.com/world
/http://hello.com
/http://hello.com/world
/www.hello.com
/www.hello.com/world
/hello.com
/hello.com/world
/hello
/hello/world

但不是这样的:

hello 
hello/world
hello.png
hello.zip
/hello/world.png
/hello/world.js

为此我需要什么正则表达式?或者,还有更好的方法? (也许通过收集一个)

【问题讨论】:

  • 为什么投反对票?似乎是一个合法的问题
  • “有更好的方法吗?”:嗯,正则表达式不能做到这一点完全健壮(由于 HTML 语言的性质)。但另一种选择是使用 HTML/XML 解析器,这对于您的简单任务可能完全是多余的。所以我会选择正则表达式。

标签: php html regex


【解决方案1】:

我猜你将“链接”定义为<a href="..."> 形式的超链接。下面的正则表达式(已经是 PHP 字符串的形式)应该是一个好的开始*:

'<\\s*a\\s*[^>]*href\\s*=\\s*"([^"]+)"'

Test this regex

当与preg_match($regex, $html, $match) 一起使用时,$match[1] 会为您提供链接,但是,它是编码形式(它可能包含 html 实体)。要删除这些,请使用html_entity_decode。

$link = html_entity_decode($match[1]);

您还应该排除只是同一站点片段的链接,即以井号开头的链接:$link[0] == '#'


*这个正则表达式不符合 HTML 语言的定义(我认为这不可能 100% 正确)。例如,对于属性未用双引号括起来的链接(它们可能未加引号或用单引号引起来),正则表达式会失败。

【讨论】:

    【解决方案2】:

    在这种情况下,PHPQuery 之类的东西可能比使用正则表达式更可取。请参阅 this answer 了解原因。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2022-10-15
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-08-20
      相关资源
      最近更新 更多