虽然您确实无法使用正则表达式可靠地 _解析_ HTML,但这不是 OP 所要求的。
相反,OP 需要一种从 HTML 文档中提取锚链接的方法,该方法可以使用正则表达式轻松且出色地处理。
上一个回复者列出的四个问题中:
- 锚的各个部分之间有多个空格
- 使用单引号而不是双引号
- 根本不使用引号来分隔 href 属性
- 具有除 href 之外的其他前导或尾随属性
只有数字 3 对 single 正则表达式解决方案提出了重大问题,但也恰好是完全不应该出现在 HTML 文档中的非标准 HTML。 (请注意,如果您发现 HTML 包含非分隔标记属性,则有一个正则表达式将匹配它们,但我认为它们不值得提取。YMMV - 您的里程可能会有所不同。)
要使用正则表达式从 HTML 中提取锚链接 (hrefs),您可以使用以下正则表达式(以注释形式):
< # a literal '<'
a # a literal 'a'
[^>]+? # one or more chars which are not '>' (non-greedy)
href= # literal 'href='
('|") # either a single or double-quote captured into group #1
([^\1]+?) # one or more chars that are not the group #1, captured into group #2
\1 # whatever capture group #1 matched
没有 cmets 是:
<a[^>]+?href=('|")([^\1]+?)\1
(请注意,我们不需要匹配任何超出最终分隔符的内容,包括标记的其余部分,因为我们只对锚链接感兴趣。)
在 JavaScript 中并假设“源”包含您希望从中提取锚链接的 HTML:
var source='<a href="double-quote test">\n'+
'<a href=\'single-quote test\'>\n'+
'<a class="foo" href="leading prop test">\n'+
'<a href="trailing prop test" class="foo">\n'+
'<a style="bar" link="baz" '+
'name="quux" '+
'href="multiple prop test" class="foo">\n'+
'<a class="foo"\n href="inline newline test"\n style="bar"\n />';
当打印到控制台时,内容如下:
<a href="double-quote test">
<a href='single-quote test'>
<a class="foo" href="leading prop test">
<a href="trailing prop test" class="foo">
<a style="bar" link="baz" name="quux" href="multiple prop test" class="foo">
<a class="foo"
href="inline newline test"
style="bar"
/>
你会写如下:
var RE=new RegExp(/<a[^>]+?href=('|")([^\1]+?)\1/gi),
match;
while(match=RE.exec(source)) {
console.log(match[2]);
}
它将以下行打印到控制台:
double-quote test
single-quote test
leading prop test
trailing prop test
multiple prop test
inline newline test
注意事项:
在 nodejs v0.5.0-pre 中测试的代码,但应该在任何现代 JavaScript 下运行。
由于正则表达式使用捕获组 #1 来标注前导分隔引号,因此生成的链接出现在捕获组 #2 中。
-
您可能希望使用以下方法验证匹配的存在、类型和长度:
if(match && typeof match === 'object' && match.length > 1) {
console.log(match[2]);
}
但它确实没有必要,因为 RegExp.exec() 在失败时返回“null”。另请注意,正确的 typeof 匹配项是“object”,而不是“Array”。