【问题标题】:PHP preg_replace match HTML attributePHP preg_replace 匹配 HTML 属性
【发布时间】:2013-03-06 16:27:25
【问题描述】:

我正在尝试从 HTML 元素中删除 title 属性。

function remove_title_attributes($input) {
    return remove_html_attribute('title', $input);
}

/**
 * To remove an attribute from an html tag
 * @param string $attr the attribute
 * @param string $str the html
 */
function remove_html_attribute($attr, $str){
    return preg_replace('/\s*'.$attr.'\s*=\s*(["\']).*?\1/', '', $str);
}

但是,它无法区分<img title="something"> 和[shortcode title="something"]。如何仅定位 HTML 标记中的代码(例如 <img> 或 <a href=""><a>)?

【问题讨论】:

  • 为此使用 HTML 解析器,而不是正则表达式函数。
  • 不要使用正则表达式解析 HTML。您无法使用正则表达式可靠地解析 HTML。一旦 HTML 与您的期望发生变化,您的代码就会被破坏。有关如何使用 PHP 模块正确解析 HTML 的示例,请参阅 htmlparsing.com/php.html。
  • 当我使用这个功能时,我没有一个完整的 HTML 文档。只是没有根标签的博客文章的正文内容。像这样的东西:<p>stuff <a href="link" title="something">linkme</a></p><p>more stuff</p><p>even more stuff</p>

标签: php html regex html-parsing


【解决方案1】:

不要使用正则表达式,而是使用 DOM 解析器。转到official reference page 并研究它。在您的情况下,您需要 DOMElement::removeAttribute() 方法。这是一个例子:

<?php

$html = '<p>stuff <a href="link" title="something">linkme</a></p><p>more stuff</p><p>even more stuff</p>';

$dom = new DOMDocument();
$dom->loadHTML($html);

$domElement = $dom->documentElement;

$a = $domElement->getElementsByTagName('a')->item(0);
$a->removeAttribute('title');

$result =  $dom->saveHTML();

【讨论】:

【解决方案2】:

我使用来自@Hast 的代码作为构建块。看起来这样可以解决问题(除非有更好的方法?)

/**
 * To remove an attribute from an html tag
 * @param string $attr the attribute
 * @param string $str the html
 */
function remove_html_attribute($attr, $input){
    //return preg_replace('/\s*'.$attr.'\s*=\s*(["\']).*?\1/', '', $input);

    $result='';

    if(!empty($input)){

        //check if the input text contains tags
        if($input!=strip_tags($input)){
            $dom = new DOMDocument();

            //use mb_convert_encoding to prevent non-ASCII characters from randomly appearing in text
            $dom->loadHTML(mb_convert_encoding($input, 'HTML-ENTITIES', 'UTF-8'));

            $domElement = $dom->documentElement;

            $taglist = array('a', 'img', 'span', 'li', 'table', 'td'); //tags to check for specified tag attribute

            foreach($taglist as $target_tag){
                $tags = $domElement->getElementsByTagName($target_tag);

                foreach($tags as $tag){
                    $tag->removeAttribute($attr);
                }
            }

            //$result =  $dom->saveHTML();
            $result = innerHTML( $domElement->firstChild ); //strip doctype/html/body tags
        }
        else{
            $result=$input;
        }
    }

    return $result; 
}

/**
 * removes the doctype/html/body tags
 */
function innerHTML($node){
  $doc = new DOMDocument();
  foreach ($node->childNodes as $child)
    $doc->appendChild($doc->importNode($child, true));

  return $doc->saveHTML();
}

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-01-06
    • 2013-11-05
    • 2016-09-23
    • 1970-01-01
    相关资源
    最近更新 更多