【发布时间】:2017-05-22 07:27:37
【问题描述】:
我需要捕获来自多个网站的所有链接。为此,我收集了整个 html 文件。我需要一个将所有这些都放在一个数组中的正则表达式。
我不想收集任何图像文件或其他代码文件。只是页面本身的 html。
我希望它收集所有这样的链接:
/https://www.hello.com
/https://www.hello.com/index.php
/https://www.hello.com/world
/https://www.hello.com/world.php
/https://www.hello.com/world.html
/https://hello.com
/https://hello.com/world
/http://www.hello.com
/http://www.hello.com/world
/http://hello.com
/http://hello.com/world
/www.hello.com
/www.hello.com/world
/hello.com
/hello.com/world
/hello
/hello/world
但不是这样的:
hello
hello/world
hello.png
hello.zip
/hello/world.png
/hello/world.js
为此我需要什么正则表达式?或者,还有更好的方法? (也许通过收集一个)
【问题讨论】:
-
为什么投反对票?似乎是一个合法的问题
-
“有更好的方法吗?”:嗯,正则表达式不能做到这一点完全健壮(由于 HTML 语言的性质)。但另一种选择是使用 HTML/XML 解析器,这对于您的简单任务可能完全是多余的。所以我会选择正则表达式。