【问题标题】:How can I remove attributes from an html tag?如何从 html 标签中删除属性?
【发布时间】:2010-10-20 16:35:25
【问题描述】:

如何使用 php 去除标签中的所有/任何属性,比如段落标签?

<p class="one" otherrandomattribute="two"><p>

【问题讨论】:

    标签: php html-parsing


    【解决方案1】:

    虽然有更好的方法,但您实际上可以使用正则表达式从 html 标记中去除参数:

    <?php
    function stripArgumentFromTags( $htmlString ) {
        $regEx = '/([^<]*<\s*[a-z](?:[0-9]|[a-z]{0,9}))(?:(?:\s*[a-z\-]{2,14}\s*=\s*(?:"[^"]*"|\'[^\']*\'))*)(\s*\/?>[^<]*)/i'; // match any start tag
    
        $chunks = preg_split($regEx, $htmlString, -1,  PREG_SPLIT_DELIM_CAPTURE);
        $chunkCount = count($chunks);
    
        $strippedString = '';
        for ($n = 1; $n < $chunkCount; $n++) {
            $strippedString .= $chunks[$n];
        }
    
        return $strippedString;
    }
    ?>
    

    上面的内容可能可以用更少的字符编写,但它可以完成工作(快速而肮脏)。

    【讨论】:

      【解决方案2】:

      使用 SimpleXML 剥离属性(PHP5 中的标准)

      <?php
      
      // define allowable tags
      $allowable_tags = '<p><a><img><ul><ol><li><table><thead><tbody><tr><th><td>';
      // define allowable attributes
      $allowable_atts = array('href','src','alt');
      
      // strip collector
      $strip_arr = array();
      
      // load XHTML with SimpleXML
      $data_sxml = simplexml_load_string('<root>'. $data_str .'</root>', 'SimpleXMLElement', LIBXML_NOERROR | LIBXML_NOXMLDECL);
      
      if ($data_sxml ) {
          // loop all elements with an attribute
          foreach ($data_sxml->xpath('descendant::*[@*]') as $tag) {
              // loop attributes
              foreach ($tag->attributes() as $name=>$value) {
                  // check for allowable attributes
                  if (!in_array($name, $allowable_atts)) {
                      // set attribute value to empty string
                      $tag->attributes()->$name = '';
                      // collect attribute patterns to be stripped
                      $strip_arr[$name] = '/ '. $name .'=""/';
                  }
              }
          }
      }
      
      // strip unallowed attributes and root tag
      $data_str = strip_tags(preg_replace($strip_arr,array(''),$data_sxml->asXML()), $allowable_tags);
      
      ?>
      

      【讨论】:

      • 这很好用,但前提是您的输入 html 格式正确 xml。否则,您必须在解析之前对输入的 html 进行一些预清理。如果您也不能完全控制原始 html 输入,那么清理这可能会非常乏味。
      【解决方案3】:

      这里有一个函数可以让你去除除你想要的属性之外的所有属性:

      function stripAttributes($s, $allowedattr = array()) {
        if (preg_match_all("/<[^>]*\\s([^>]*)\\/*>/msiU", $s, $res, PREG_SET_ORDER)) {
         foreach ($res as $r) {
           $tag = $r[0];
           $attrs = array();
           preg_match_all("/\\s.*=(['\"]).*\\1/msiU", " " . $r[1], $split, PREG_SET_ORDER);
           foreach ($split as $spl) {
            $attrs[] = $spl[0];
           }
           $newattrs = array();
           foreach ($attrs as $a) {
            $tmp = explode("=", $a);
            if (trim($a) != "" && (!isset($tmp[1]) || (trim($tmp[0]) != "" && !in_array(strtolower(trim($tmp[0])), $allowedattr)))) {
      
            } else {
                $newattrs[] = $a;
            }
           }
           $attrs = implode(" ", $newattrs);
           $rpl = str_replace($r[1], $attrs, $tag);
           $s = str_replace($tag, $rpl, $s);
         }
        }
        return $s;
      }
      

      例如:

      echo stripAttributes('<p class="one" otherrandomattribute="two">');
      

      或者,如果你例如。想要保留“class”属性:

      echo stripAttributes('<p class="one" otherrandomattribute="two">', array('class'));
      

      或者

      假设您要向收件箱发送消息并且您使用 CKEDITOR 编写消息,您可以按如下方式分配函数并在发送前将其回显到 $message 变量。请注意,名为 stripAttributes() 的函数将删除所有不必要的 html 标记。我试过了,效果很好。我只看到了我添加的格式,比如粗体等。

      $message = stripAttributes($_POST['message']);
      

      或 您可以echo $message;进行预览。

      【讨论】:

        【解决方案4】:

        HTML Purifier 是使用 PHP 清理 HTML 的更好工具之一。

        【讨论】:

          【解决方案5】:

          老实说,我认为唯一明智的方法是使用标签 属性白名单与HTML Purifier 库。此处的示例脚本:

          <html><body>
          
          <?php
          
          require_once '../includes/htmlpurifier-4.5.0-lite/library/HTMLPurifier/Bootstrap.php';
          spl_autoload_register(array('HTMLPurifier_Bootstrap', 'autoload'));
          
          $config = HTMLPurifier_Config::createDefault();
          $config->set('HTML.Allowed', 'p,b,a[href],i,br,img[src]');
          $config->set('URI.Base', 'http://www.example.com');
          $config->set('URI.MakeAbsolute', true);
          
          $purifier = new HTMLPurifier($config);
          
          $dirty_html = "
            <a href=\"http://www.google.de\">broken a href link</a
            fnord
          
            <x>y</z>
            <b>c</p>
            <script>alert(\"foo!\");</script>
          
            <a href=\"javascript:alert(history.length)\">Anzahl besuchter Seiten</a>
            <img src=\"www.example.com/bla.gif\" />
            <a href=\"http://www.google.de\">missing end tag
           ende 
          ";
          
          $clean_html = $purifier->purify($dirty_html);
          
          print "<h1>dirty</h1>";
          print "<pre>" . htmlentities($dirty_html) . "</pre>";
          
          print "<h1>clean</h1>";
          print "<pre>" . htmlentities($clean_html) . "</pre>";
          
          ?>
          
          </body></html>
          

          这会产生以下干净、符合标准的 HTML 片段:

          <a href="http://www.google.de">broken a href link</a>fnord
          
          y
          <b>c
          <a>Anzahl besuchter Seiten</a>
          <img src="http://www.example.com/www.example.com/bla.gif" alt="bla.gif" /><a href="http://www.google.de">missing end tag
          ende 
          </a></b>
          

          在您的情况下,白名单将是:

          $config->set('HTML.Allowed', 'p');
          

          【讨论】:

            【解决方案6】:

            您还可以研究 html 净化器。诚然,它非常臃肿,如果仅考虑这个特定示例,它可能不适合您的需求,但它或多或少地提供了对可能的敌对 html 的“防弹”净化。您还可以选择允许或禁止某些属性(高度可配置)。

            http://htmlpurifier.org/

            【讨论】:

              猜你喜欢
              • 2012-02-17
              • 2011-03-02
              • 2014-09-27
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              相关资源
              最近更新 更多