【发布时间】:2013-03-06 16:27:25
【问题描述】:
我正在尝试从 HTML 元素中删除 title 属性。
function remove_title_attributes($input) {
return remove_html_attribute('title', $input);
}
/**
* To remove an attribute from an html tag
* @param string $attr the attribute
* @param string $str the html
*/
function remove_html_attribute($attr, $str){
return preg_replace('/\s*'.$attr.'\s*=\s*(["\']).*?\1/', '', $str);
}
但是,它无法区分<img title="something"> 和[shortcode title="something"]。如何仅定位 HTML 标记中的代码(例如 <img> 或 <a href=""><a>)?
【问题讨论】:
-
为此使用 HTML 解析器,而不是正则表达式函数。
-
不要使用正则表达式解析 HTML。您无法使用正则表达式可靠地解析 HTML。一旦 HTML 与您的期望发生变化,您的代码就会被破坏。有关如何使用 PHP 模块正确解析 HTML 的示例,请参阅 htmlparsing.com/php.html。
-
当我使用这个功能时,我没有一个完整的 HTML 文档。只是没有根标签的博客文章的正文内容。像这样的东西:
<p>stuff <a href="link" title="something">linkme</a></p><p>more stuff</p><p>even more stuff</p>
标签: php html regex html-parsing