您不需要为此使用服务器端 HTML 解析器,这在 imo 中完全是多余的。以下是一个显然可能被某些 HTML 结构破坏的示例,但对于大多数标记来说,它不会有任何问题 - 并且比服务器端 HTML 解析器更优化。
$html = '
<h2>My text one</h2>
<h2>My text two</h2>
text ...
<h2>My text three</h2>
';
preg_match_all
/// the following preg match will find all <h2> mark-up, even if
/// the content of the h2 splits over new lines - due to the `s` switch
/// It is a non-greedy match too - thanks to the `.+?` so it shouldn't
/// have problems with spanning over more than one h2 tag. It will only
/// really break down if you have a h2 as a descendant of a h2 - which
/// would be illegal html - or if you have a `>` in one of your h2's
/// attributes i.e. <h2 title="this>will break">Text</h2> which again
/// is illegal as they should be encoded.
preg_match_all(
'#(<)h2([^>]*>.+?</)h2(>)#is',
$html,
$matches,
PREG_OFFSET_CAPTURE|PREG_SET_ORDER
);
替换和重建
/// Because you wanted to only replace the 2nd item use the following.
/// You could however make this code as general or as specific as you wanted.
/// The following works because the surrounding content for the found
/// $matches was stored using the grouping brackets in the regular
/// expression. This means you could easily change the regexp, and the
/// following code would still work.
/// to get a better understanding of what is going on it would be best
/// to `echo '<xmp>';print_r( $matches );echo '/<xmp>';`
if ( isset($matches[1][0]) ) {
$html = substr( $html, 0, $matches[1][0][1] ) .
$matches[1][1][0] . 'h3' .
$matches[1][2][0] . 'h3' .
$matches[1][3][0] .
substr( $html, $matches[1][0][1] + strlen($matches[1][0][0]) );
}
我不知道为什么很多人说要使用客户端 JavaScript 来进行这种更改,PHP 代表 PHP: Hypertext Preprocessor,它旨在预处理超文本。 OP 只提到过 PHP 函数,并用 PHP 标记了这篇文章,所以没有任何东西指向客户端。
确实,尽管可以并且应该尽可能使用客户端来减轻服务器端的处理,但不建议将其用于诸如标题之类的核心结构标签 - 屏幕阅读器和搜索引擎机器人将依赖它。最好使用客户端 JavaScript 来增强用户体验。如果您使用它来极大地增强您网站的功能,您最好确保您的整个用户群都支持它。
但是,如果你们中的任何人提到 Node.js 和 JSDOM,我会很高兴地同意。