【问题标题】:How can I convert HTML character references (ף) to regular UTF-8?如何将 HTML 字符引用 (ף) 转换为常规 UTF-8?
【发布时间】:2011-04-03 16:52:24
【问题描述】:

我有一些包含字符引用的希伯来语网站,例如:נוף

如果我将文件保存为 .html 并以 UTF-8 编码查看,我只能查看这些字母。

如果我尝试将其作为常规文本文件打开,则 UTF-8 编码不会显示正确的输出。

我注意到,如果我打开一个文本编辑器并用 UTF-8 编写希伯来语,在此示例中每个字符占用两个字节而不是 4 个字节行 (ו)

如果这是 UTF-16 或任何其他类型的 UTF 字母表示有什么想法吗?

如果可能,如何将其转换为普通字母?

使用最新的 PHP 版本。

【问题讨论】:

  • 现在您的文件中有什么:נוף 或 נוף?
  • 我想知道如何将其转换为常规的 utf-8 并且我想知道这些字符是什么?这是 utf-16 的表示还是别的什么?

标签: php html utf-8 character-reference


【解决方案1】:

character references 通过以十进制 (&amp;#<i>n</i>;) 或十六进制 (&amp;#x<i>n</i>;) 表示法指定该字符的代码点来引用 ISO 10646 中的字符。

您可以使用html_entity_decode 解码此类字符引用以及entities defined for HTML 4 的实体引用,因此其他引用(如&amp;lt;、&amp;gt;、&amp;amp;)也将被解码:

$str = html_entity_decode($str, ENT_NOQUOTES, 'UTF-8');

如果你只想解码数字字符引用,你可以使用这个:

function html_dereference($match) {
    if (strtolower($match[1][0]) === 'x') {
        $codepoint = intval(substr($match[1], 1), 16);
    } else {
        $codepoint = intval($match[1], 10);
    }
    return mb_convert_encoding(pack('N', $codepoint), 'UTF-8', 'UTF-32BE');
}
$str = preg_replace_callback('/&#(x[0-9a-f]+|[0-9]+);/i', 'html_dereference', $str);

正如YuriKolovsky 和thirtydot 在another question 中指出的那样,似乎浏览器供应商确实“默默地”就字符引用映射达成了一致,这确实与规范不同,并且没有记录。

似乎有一些字符引用通常会映射到Latin 1 supplement,但实际上映射到不同的字符。这是由于映射宁愿从 Windows-1252 而不是 ISO 8859-1 中映射字符产生的映射,Unicode 字符集建立在 ISO 8859-1 之上。 Jukka Korpela 写了一个extensive article on this topic。

现在这里是处理这个怪癖的上述函数的扩展:

function html_character_reference_decode($string, $encoding='UTF-8', $fixMappingBug=true) {
    $deref = function($match) use ($encoding, $fixMappingBug) {
        if (strtolower($match[1][0]) === "x") {
            $codepoint = intval(substr($match[1], 1), 16);
        } else {
            $codepoint = intval($match[1], 10);
        }
        // @see http://www.cs.tut.fi/~jkorpela/www/windows-chars.html
        if ($fixMappingBug && $codepoint >= 130 && $codepoint <= 159) {
            $mapping = array(
                8218, 402, 8222, 8230, 8224, 8225, 710, 8240, 352, 8249,
                338, 141, 142, 143, 144, 8216, 8217, 8220, 8221, 8226,
                8211, 8212, 732, 8482, 353, 8250, 339, 157, 158, 376);
            $codepoint = $mapping[$codepoint-130];
        }
        return mb_convert_encoding(pack("N", $codepoint), $encoding, "UTF-32BE");
    };
    return preg_replace_callback('/&#(x[0-9a-f]+|[0-9]+);/i', $deref, $string);
}

如果anonymous functions 不可用(5.3.0 引入),您也可以使用create_function:

$deref = create_function('$match', '
    $encoding = '.var_export($encoding, true).';
    $fixMappingBug = '.var_export($fixMappingBug, true).';
    if (strtolower($match[1][0]) === "x") {
        $codepoint = intval(substr($match[1], 1), 16);
    } else {
        $codepoint = intval($match[1], 10);
    }
    // @see http://www.cs.tut.fi/~jkorpela/www/windows-chars.html
    if ($fixMappingBug && $codepoint >= 130 && $codepoint <= 159) {
        $mapping = array(
            8218, 402, 8222, 8230, 8224, 8225, 710, 8240, 352, 8249,
            338, 141, 142, 143, 144, 8216, 8217, 8220, 8221, 8226,
            8211, 8212, 732, 8482, 353, 8250, 339, 157, 158, 376);
        $codepoint = $mapping[$codepoint-130];
    }
    return mb_convert_encoding(pack("N", $codepoint), $encoding, "UTF-32BE");
');

这是另一个尝试遵守 behavior of HTML 5 的函数:

function html5_decode($string, $flags=ENT_COMPAT, $charset='UTF-8') {
    $deref = function($match) use ($flags, $charset) {
        if ($match[1][0] === '#') {
            if (strtolower($match[1][0]) === '#') {
                $codepoint = intval(substr($match[1], 2), 16);
            } else {
                $codepoint = intval(substr($match[1], 1), 10);
            }

            // HTML 5 specific behavior
            // @see http://dev.w3.org/html5/spec/tokenization.html#tokenizing-character-references

            // handle Windows-1252 mismapping
            // @see http://www.cs.tut.fi/~jkorpela/www/windows-chars.html
            // @see http://dev.w3.org/html5/spec/tokenization.html#table-charref-overrides
            $overrides = array(
                0x00=>0xFFFD,0x80=>0x20AC,0x82=>0x201A,0x83=>0x0192,0x84=>0x201E,
                0x85=>0x2026,0x86=>0x2020,0x87=>0x2021,0x88=>0x02C6,0x89=>0x2030,
                0x8A=>0x0160,0x8B=>0x2039,0x8C=>0x0152,0x8E=>0x017D,0x91=>0x2018,
                0x92=>0x2019,0x93=>0x201C,0x94=>0x201D,0x95=>0x2022,0x96=>0x2013,
                0x97=>0x2014,0x98=>0x02DC,0x99=>0x2122,0x9A=>0x0161,0x9B=>0x203A,
                0x9C=>0x0153,0x9E=>0x017E,0x9F=>0x0178);
            if (isset($windows1252Mapping[$codepoint])) {
                $codepoint = $windows1252Mapping[$codepoint];
            }

            if (($codepoint >= 0xD800 && $codepoint <= 0xDFFF) || $codepoint > 0x10FFFF) {
                $codepoint = 0xFFFD;
            }
            if (($codepoint >= 0x0001 && $codepoint <= 0x0008) ||
                ($codepoint >= 0x000E && $codepoint <= 0x001F) ||
                ($codepoint >= 0x007F && $codepoint <= 0x009F) ||
                ($codepoint >= 0xFDD0 && $codepoint <= 0xFDEF) ||
                in_array($codepoint, array(
                    0x000B, 0xFFFE, 0xFFFF, 0x1FFFE, 0x1FFFF, 0x2FFFE, 0x2FFFF,
                    0x3FFFE, 0x3FFFF, 0x4FFFE, 0x4FFFF, 0x5FFFE, 0x5FFFF, 0x6FFFE,
                    0x6FFFF, 0x7FFFE, 0x7FFFF, 0x8FFFE, 0x8FFFF, 0x9FFFE, 0x9FFFF,
                    0xAFFFE, 0xAFFFF, 0xBFFFE, 0xBFFFF, 0xCFFFE, 0xCFFFF, 0xDFFFE,
                    0xDFFFF, 0xEFFFE, 0xEFFFF, 0xFFFFE, 0xFFFFF, 0x10FFFE, 0x10FFFF))) {
                $codepoint = 0xFFFD;
            }
            return mb_convert_encoding(pack("N", $codepoint), $charset, "UTF-32BE");
        } else {
            return html_entity_decode($match[0], $flags, $charset);
        }   
    };
    return preg_replace_callback('/&(#(?:x[0-9a-f]+|[0-9]+)|[A-Za-z0-9]+);/i', $deref, $string);
}

我还注意到,在 PHP 5.4.0 中,html_entity_decode function 添加了另一个名为 ENT_HTML5 的标志,用于 HTML 5 行为。

【讨论】:

  • 使用mb_convert_encoding而不是iconv有什么特别的原因吗?
  • 你可以将一个本地字符集表示的字符串转换成另一个字符集表示的字符串,可能是Unicode字符集。支持的字符集取决于系统的 iconv 实现。
  • 很公平。还不错,我只是对选择更好奇...+1
  • 像 — 这样的 Microsoft Windows 字符引用呢? ;p
  • Alohci pointed out to me 表示“字符覆盖映射在此处的 HTML5 规范中正式指定:http://dev.w3.org/html5/spec/tokenization.html#table-charref-overrides”。您更新的功能是否匹配?
【解决方案2】:

这些是 XML Character References。你想用html_entity_decode()解码它们:

$string = html_entity_decode($string, ENT_QUOTES, 'UTF-8');

有关详细信息,您可以在 Google 上搜索相关实体。请参阅以下几个示例:

  1. Hebrew Characters
  2. HTML Entities for Hebrew Characters
  3. UTF-8 Encoding Table with HTML entities

【讨论】:

  • 那些是不是实体,甚至不是实体引用。这些只是字符引用。
  • @Gumbo:很公平。他们没有使用命名实体......但概念几乎相同(除了不需要地图)。我将编辑答案以反映...
猜你喜欢
  • 2014-07-03
  • 2015-04-04
  • 2023-03-24
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-04-03
  • 2011-06-25
相关资源
最近更新 更多