【问题标题】:How do I read UTF-8 characters via a pointer?如何通过指针读取 UTF-8 字符?
【发布时间】:2011-02-26 06:12:59
【问题描述】:

假设我在内存中存储了 UTF-8 内容,如何使用指针读取字符?我想我需要注意指示多字节字符的第 8 位,但我究竟如何将序列转换为有效的 Unicode 字符?另外,wchar_t 是存储单个 Unicode 字符的正确类型吗?

这就是我的想法:

wchar_t readNextChar (char*& p) { wchar_t unicodeChar; char ch = *p++; if ((ch & 128) != 0) { // This is a multi-byte character, what do I do now? // char chNext = *p++; // ... but how do I assemble the Unicode character? ... } ... return unicodeChar; }

【问题讨论】:

  • 说“Unicode 字符的宽度”是没有意义的。您需要适应编码。根据您的平台,wchar_t 的大小可能不同。在类 Unix 操作系统上,它通常是 32 位的,因此您可以在其中存储 UTF-32 编码的 Unicode 字符,在 Windows 上它是 16 位的,因此它可以采用 UTF-16 编码的 Unicode 字符。
  • 除了返回宽字符之外,您的readNextChar 函数还必须提供信息以正确更新p。 UTF-8(以及 UTF-16,就此而言)是可变长度编码,您的调用者不能假定指针中的常量或简单增量。

标签: c++ unicode utf-8 character-encoding


【解决方案1】:

您必须将 UTF-8 位模式解码为其未编码的 UTF-32 表示。如果您想要实际的 Unicode 代码点,则必须使用 32 位数据类型。

在 Windows 上,wchar_t 不够大,因为它只有 16 位。您必须改用unsigned int 或unsigned long。仅在处理 UTF-16 代码单元时才使用 wchar_t。

在其他平台上,wchar_t 通常是 32 位的。但是在编写可移植代码时,您应该远离wchar_t,除非绝对需要(例如std::wstring)。

试试这样的:

#define IS_IN_RANGE(c, f, l)    (((c) >= (f)) && ((c) <= (l)))

u_long readNextChar (char* &p) 
{  
    // TODO: since UTF-8 is a variable-length
    // encoding, you should pass in the input
    // buffer's actual byte length so that you
    // can determine if a malformed UTF-8
    // sequence would exceed the end of the buffer...

    u_char c1, c2, *ptr = (u_char*) p;
    u_long uc = 0;
    int seqlen;
    // int datalen = ... available length of p ...;    

    /*
    if( datalen < 1 )
    {
        // malformed data, do something !!!
        return (u_long) -1;
    }
    */

    c1 = ptr[0];

    if( (c1 & 0x80) == 0 )
    {
        uc = (u_long) (c1 & 0x7F);
        seqlen = 1;
    }
    else if( (c1 & 0xE0) == 0xC0 )
    {
        uc = (u_long) (c1 & 0x1F);
        seqlen = 2;
    }
    else if( (c1 & 0xF0) == 0xE0 )
    {
        uc = (u_long) (c1 & 0x0F);
        seqlen = 3;
    }
    else if( (c1 & 0xF8) == 0xF0 )
    {
        uc = (u_long) (c1 & 0x07);
        seqlen = 4;
    }
    else
    {
        // malformed data, do something !!!
        return (u_long) -1;
    }

    /*
    if( seqlen > datalen )
    {
        // malformed data, do something !!!
        return (u_long) -1;
    }
    */

    for(int i = 1; i < seqlen; ++i)
    {
        c1 = ptr[i];

        if( (c1 & 0xC0) != 0x80 )
        {
            // malformed data, do something !!!
            return (u_long) -1;
        }
    }

    switch( seqlen )
    {
        case 2:
        {
            c1 = ptr[0];

            if( !IS_IN_RANGE(c1, 0xC2, 0xDF) )
            {
                // malformed data, do something !!!
                return (u_long) -1;
            }

            break;
        }

        case 3:
        {
            c1 = ptr[0];
            c2 = ptr[1];

            switch (c1)
            {
                case 0xE0:
                    if (!IS_IN_RANGE(c2, 0xA0, 0xBF))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;

                case 0xED:
                    if (!IS_IN_RANGE(c2, 0x80, 0x9F))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;

                default:
                    if (!IS_IN_RANGE(c1, 0xE1, 0xEC) && !IS_IN_RANGE(c1, 0xEE, 0xEF))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;
            }

            break;
        }

        case 4:
        {
            c1 = ptr[0];
            c2 = ptr[1];

            switch (c1)
            {
                case 0xF0:
                    if (!IS_IN_RANGE(c2, 0x90, 0xBF))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;

                case 0xF4:
                    if (!IS_IN_RANGE(c2, 0x80, 0x8F))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;

                default:
                    if (!IS_IN_RANGE(c1, 0xF1, 0xF3))
                    {
                        // malformed data, do something !!!
                        return (u_long) -1;
                    }
                    break;                
            }

            break;
        }
}

    for(int i = 1; i < seqlen; ++i)
    {
        uc = ((uc << 6) | (u_long)(ptr[i] & 0x3F));
    }

    p += seqlen;
    return uc; 
}

【讨论】:

  • @Remy & @Jen:wchar_t 的确切宽度没有定义。在 GCC 上(至少在 Linux 上)wchar_t 是 32 位的,因此在不进行多字节编码的情况下保存 Unicode 字符当然就足够了。
  • 我刚刚使用 g++ 3.3.4 编译了这段代码,给我留下了深刻的印象:编译器将所有代码从大的 switch 语句移到了设置 seqlen 变量的位置。也许这对原始代码也有好处,变得更具可读性。
  • @sbi:在 MSVC 中它是 16 位的,而 Windows API 需要 16 位,在 Windows 编程中几乎强制使用 16 位。
  • @DeadMG:我一定是错过了你从这个问题中得出这些结论的心理能力。
  • @DavidHaim: u8"?" 被编码为字节F0 9F 98 80,这是U+1F600 GRINNING FACE 的正确UTF-8 序列。我已经调整了代码以正确处理。这只是 IS_IN_RANGE() 应用方式的逻辑错误。
【解决方案2】:

这是一个计算 UTF-8 字节数的快速宏

#define UTF8_CHAR_LEN( byte ) (( 0xE5000000 >> (( byte >> 3 ) & 0x1e )) & 3 ) + 1

这将帮助您检测 UTF-8 字符的大小以便于解析。

【讨论】:

    【解决方案3】:

    如果您需要解码 UTF-8,您需要开发一个 UTF-8 解析器。 UTF-8 是一种可变长度编码(1 到 4 个字节),因此您确实必须编写一个符合标准的解析器:例如,参见 wikipedia。

    如果您不想编写自己的解析器,我建议使用库。例如,您会在 glib 中发现(我个人使用过 Glib::ustring,glib 的 C++ 包装器)以及任何好的通用库中。

    编辑:

    我认为 C++0x 也会包含 UTF-8 支持,但我不是专家...

    my2c

    【讨论】:

      【解决方案4】:

      另外,wchar_t 是否适合存储单个 Unicode 字符?

      在 Linux 上,是的。在 Windows 上,wchar_t 表示 UTF-16 代码单元,不一定是字符。

      即将推出的 C++0x 标准将提供 char16_t 和 char32_t 类型,旨在表示 UTF-16 和 UTF-32。

      如果在char32_t 不可用且wchar_t 不足的系统上,请使用uint32_t 存储Unicode 字符。

      【讨论】:

        【解决方案5】:

        这是我在纯 ANSI-C 中的解决方案,包括针对极端情况的单元测试。

        注意int 必须至少为 32 位宽。否则你必须改变codepoint的定义。

        #include <assert.h>
        #include <errno.h>
        #include <stdio.h>
        #include <stdlib.h>
        
        typedef unsigned char byte;
        typedef unsigned int codepoint;
        
        /**
         * Reads the next UTF-8-encoded character from the byte array ranging
         * from {@code *pstart} up to, but not including, {@code end}. If the
         * conversion succeeds, the {@code *pstart} iterator is advanced,
         * the codepoint is stored into {@code *pcp}, and the function returns
         * 0. Otherwise the conversion fails, {@code errno} is set to
         * {@code EILSEQ} and the function returns -1.
         */
        int
        from_utf8(const byte **pstart, const byte *end, codepoint *pcp) {
                size_t len, i;
                codepoint cp, min;
                const byte *buf;
        
                buf = *pstart;
                if (buf == end)
                        goto error;
        
                if (buf[0] < 0x80) {
                        len = 1;
                        min = 0;
                        cp = buf[0];
                } else if (buf[0] < 0xC0) {
                        goto error;
                } else if (buf[0] < 0xE0) {
                        len = 2;
                        min = 1 << 7;
                        cp = buf[0] & 0x1F;
                } else if (buf[0] < 0xF0) {
                        len = 3;
                        min = 1 << (5 + 6);
                        cp = buf[0] & 0x0F;
                } else if (buf[0] < 0xF8) {
                        len = 4;
                        min = 1 << (4 + 6 + 6);
                        cp = buf[0] & 0x07;
                } else {
                        goto error;
                }
        
                if (buf + len > end)
                        goto error;
        
                for (i = 1; i < len; i++) {
                        if ((buf[i] & 0xC0) != 0x80)
                                goto error;
                        cp = (cp << 6) | (buf[i] & 0x3F);
                }
        
                if (cp < min)
                        goto error;
        
                if (0xD800 <= cp && cp <= 0xDFFF)
                        goto error;
        
                if (0x110000 <= cp)
                        goto error;
        
                *pstart += len;
                *pcp = cp;
                return 0;
        
        error:
                errno = EILSEQ;
                return -1;
        }
        
        static void
        assert_valid(const byte **buf, const byte *end, codepoint expected) {
                codepoint cp;
        
                if (from_utf8(buf, end, &cp) == -1) {
                        fprintf(stderr, "invalid unicode sequence for codepoint %u\n", expected);
                        exit(EXIT_FAILURE);
                }
        
                if (cp != expected) {
                        fprintf(stderr, "expected %u, got %u\n", expected, cp);
                        exit(EXIT_FAILURE);
                }
        }
        
        static void
        assert_invalid(const char *name, const byte **buf, const byte *end) {
                const byte *p;
                codepoint cp;
        
                p = *buf + 1;
                if (from_utf8(&p, end, &cp) == 0) {
                        fprintf(stderr, "unicode sequence \"%s\" unexpectedly converts to %#x.\n", name, cp);
                        exit(EXIT_FAILURE);
                }
                *buf += (*buf)[0] + 1;
        }
        
        static const byte valid[] = {
                0x00, /* first ASCII */
                0x7F, /* last ASCII */
                0xC2, 0x80, /* first two-byte */
                0xDF, 0xBF, /* last two-byte */
                0xE0, 0xA0, 0x80, /* first three-byte */
                0xED, 0x9F, 0xBF, /* last before surrogates */
                0xEE, 0x80, 0x80, /* first after surrogates */
                0xEF, 0xBF, 0xBF, /* last three-byte */
                0xF0, 0x90, 0x80, 0x80, /* first four-byte */
                0xF4, 0x8F, 0xBF, 0xBF /* last codepoint */
        };
        
        static const byte invalid[] = {
                1, 0x80,
                1, 0xC0,
                1, 0xC1,
                2, 0xC0, 0x80,
                2, 0xC2, 0x00,
                2, 0xC2, 0x7F,
                2, 0xC2, 0xC0,
                3, 0xE0, 0x80, 0x80,
                3, 0xE0, 0x9F, 0xBF,
                3, 0xED, 0xA0, 0x80,
                3, 0xED, 0xBF, 0xBF,
                4, 0xF0, 0x80, 0x80, 0x80,
                4, 0xF0, 0x8F, 0xBF, 0xBF,
                4, 0xF4, 0x90, 0x80, 0x80
        };
        
        int
        main() {
                const byte *p, *end;
        
                p = valid;
                end = valid + sizeof valid;
                assert_valid(&p, end, 0x000000);
                assert_valid(&p, end, 0x00007F);
                assert_valid(&p, end, 0x000080);
                assert_valid(&p, end, 0x0007FF);
                assert_valid(&p, end, 0x000800);
                assert_valid(&p, end, 0x00D7FF);
                assert_valid(&p, end, 0x00E000);
                assert_valid(&p, end, 0x00FFFF);
                assert_valid(&p, end, 0x010000);
                assert_valid(&p, end, 0x10FFFF);
        
                p = invalid;
                end = invalid + sizeof invalid;
                assert_invalid("80", &p, end);
                assert_invalid("C0", &p, end);
                assert_invalid("C1", &p, end);
                assert_invalid("C0 80", &p, end);
                assert_invalid("C2 00", &p, end);
                assert_invalid("C2 7F", &p, end);
                assert_invalid("C2 C0", &p, end);
                assert_invalid("E0 80 80", &p, end);
                assert_invalid("E0 9F BF", &p, end);
                assert_invalid("ED A0 80", &p, end);
                assert_invalid("ED BF BF", &p, end);
                assert_invalid("F0 80 80 80", &p, end);
                assert_invalid("F0 8F BF BF", &p, end);
                assert_invalid("F4 90 80 80", &p, end);
        
                return 0;
        }
        

        【讨论】:

        • 史诗般的失败。海报是用 C++ 写的,你违反了这么多 C++ 习语,我都数不过来了。
        • 所以我会为你数数。 (1) 我包含了 C 头文件而不是 C++ 头文件。 (2) 我使用指针而不是引用。 (3) 我没有使用命名空间,而是声明了我的函数static。 (4) 我用函数范围声明了循环变量。但另一方面,我并没有发明奇怪的类型名称(u_long、u_char)并不一致地使用它们(u_char vs. uchar)并且没有声明它们。我还设法完全避免了任何类型转换(公认的答案使用了很多,这也是 C 风格。)
        • 顺便说一句,在这种情况下,我使用指针而不是引用是完全有意的,因为 pstart 和 end 在调用方看来会非常相似。 from_utf8(start, end, &amp;cp)。谁能猜到start 被修改而end 没有被修改?
        • @Roland:呃,通过查看函数的签名? @DeadMG:公平地说,这是一个 C 解决方案。
        • @DeadMG:至少我不会抱怨你的 Lua 解决方案是糟糕的 C++ 代码。 (不过,我会抱怨我的 C++ 编译器无法编译它。)
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2017-05-04
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2015-07-26
        相关资源
        最近更新 更多