【发布时间】:2015-12-04 17:06:20
【问题描述】:
我正在尝试从包含 HTML 的数据库列中提取包含 www.domain.com 的 url。正则表达式必须过滤掉 www2.domain.com 实例和外部 URL,如 www.domainxyz.com。它应该只搜索正确编码的锚链接。
这是我目前所拥有的:
<?php
$content = '<html>
<title>Random Website</title>
<body>
Click <a href="http://domainxyz.com">here</a> for foobar
Another site is http://www.domain.com
<a href="http://www.domain.com/test">Test 1</a>
<a href="http://www2.domain.com/test">Test 2</a>
<Strong>NOT A LINK</strong>
</body>
</html>';
$regex = "((https?)\:\/\/)?";
$regex .= "([a-z0-9-.]*)\.([a-z]{2,4})";
$regex .= "(\/([a-z0-9+\$_-]\.?)+)*\/?";
$regex .= "(\?[a-z+&\$_.-][a-z0-9;:@&%=+\/\$_.-]*)?";
$regex .= "(#[a-z_.-][a-z0-9+\$_.-]*)?";
$regex .= "([www\.domain\.com])";
$matches = array(); //create array
$pattern = "/$regex/";
preg_match_all($pattern, $content, $matches);
print_r(array_values(array_unique($matches[0])));
echo "<br><br>";
echo implode("<br>", array_values(array_unique($matches[0])));
?>
我正在寻找这个仅查找和输出http://www.domain.com/test。
如何修改我的正则表达式来完成此操作?
【问题讨论】:
-
基于 DOMDocument 和 DOMXPath 的解决方案怎么样?我看你只是提取 href 属性值,对吧?
-
谢谢,我考虑过这个,但如果从数据库查询中获取 html,这样的解决方案可行吗?
-
请查看this code。我建议在这里只使用正则表达式作为最后的手段。