【问题标题】:Regex - Convert HTML to valid XML tag [duplicate]正则表达式 - 将 HTML 转换为有效的 XML 标记 [重复]
【发布时间】:2012-06-07 17:32:51
【问题描述】:

我需要帮助编写一个将 HTML 字符串转换为有效 XML 标记名称的正则表达式函数。例如:它需要一个字符串并执行以下操作:

  • 如果字符串中出现字母或下划线,则保留它
  • 如果出现任何其他字符,则会将其从输出字符串中删除。
  • 如果单词或字母之间出现任何其他字符,则将其替换为下划线。
Ex:
Input: Date Created
Ouput: Date_Created

Input: Date<br/>Created
Output: Date_Created

Input: Date\nCreated
Output: Date_Created

Input: Date    1 2 3 Created
Output: Date_Created

基本上,正则表达式函数应该将 HTML 字符串转换为有效的 XML 标记。

【问题讨论】:

  • 您的问题是“我想写”,但它读起来就像一个需求列表,正在等待某人删除所需的魔法正则表达式代码。无论如何都不清楚您认为 XML 标记是什么,输出示例不包含任何内容。
  • @JackManey:现在有超过 4000 人支持..?嘘。
  • 如果这种情况千百年来只出现一次,那有什么问题,只是在你的测试代码中快速添加一个quick and dirty patch-up!并使用正则表达式代替 DOM...

标签: php regex


【解决方案1】:

一点正则表达式和一点标准函数:

function mystrip($s)
{
        // add spaces around angle brackets to separate tag-like parts
        // e.g. "<br />" becomes " <br /> "
        // then let strip_tags take care of removing html tags
        $s = strip_tags(str_replace(array('<', '>'), array(' <', '> '), $s));

        // any sequence of characters that are not alphabet or underscore
        // gets replaced by a single underscore
        return preg_replace('/[^a-z_]+/i', '_', $s);
}

【讨论】:

    【解决方案2】:

    我相信以下应该可行。

    preg_replace('/[^A-Za-z_]+(.*)?([^A-Za-z_]+)?/', '_', $string);
    

    正则表达式[^A-Za-z_]+ 的第一部分匹配一个或多个非字母或下划线字符。正则表达式的结尾部分是相同的,只是它是可选的。这是为了允许中间部分 (.*)?(也是可选的)捕获两个列入黑名单的字符之间的任何字符(甚至是字母和下划线)。

    【讨论】:

      【解决方案3】:

      应该可以使用:

      $text = preg_replace( '/(?<=[a-zA-Z])[^a-zA-Z_]+(?=[a-zA-Z])/', '_', $text );
      

      因此,可以通过环视来查看前后是否有一个字母字符,并替换它之间的任何非字母/非下划线。

      【讨论】:

        【解决方案4】:

        试试这个

        $result = preg_replace('/([\d\s]|<[^<>]+>)/', '_', $subject);
        

        说明

        "
        (               # Match the regular expression below and capture its match into backreference number 1
                           # Match either the regular expression below (attempting the next alternative only if this one fails)
              [\d\s]          # Match a single character present in the list below
                                 # A single digit 0..9
                                 # A whitespace character (spaces, tabs, and line breaks)
           |               # Or match regular expression number 2 below (the entire group fails if this one fails to match)
              <               # Match the character “<” literally
              [^<>]           # Match a single character NOT present in the list “<>”
                 +               # Between one and unlimited times, as many times as possible, giving back as needed (greedy)
              >               # Match the character “>” literally
        )
        "
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2023-03-11
          • 1970-01-01
          • 2014-04-02
          • 2011-11-21
          • 2013-03-18
          • 2021-04-24
          • 2016-03-25
          相关资源
          最近更新 更多