【问题标题】:C++ string of greek characters and .at() operator希腊字符的 C++ 字符串和 .at() 运算符
【发布时间】:2018-03-12 01:02:25
【问题描述】:

使用英文字符很容易从字符串中提取一个字符,例如,下面的代码应该有 y 作为输出:

string my_word;
cout << my_word.at(1);

如果我尝试对希腊字符做同样的事情,我会得到一个有趣的字符:

string my_word = "λογος";
cout << my_word.at(1);

输出:

�

我的问题是:我该怎么做才能使 .at() 或任何类似的功能起作用?

非常感谢!

【问题讨论】:

标签: c++


【解决方案1】:

std::string 是一个窄字符序列char。但是在使用 utf-8 语言环境时,许多国家字母表使用不止一个字符来编码单个字母。因此,当您使用s.at(0) 时,您会得到整个字母的一半甚至更少。您应该使用宽字符:std::wstring 而不是 std::string、std::wcout 而不是 std::cout 和 L"λογος" 作为字符串文字。

此外,您应该在使用 std::locale 进行任何打印之前设置正确的区域设置。

本案例的代码示例:

#include <iostream>
#include <string>
#include <locale>

int main(int, char**) {
    std::locale::global(std::locale("en_US.utf8"));
    std::wcout.imbue(std::locale());
    std::wstring s = L"λογος";
    std::wcout << s.at(0) << std::endl;
    return 0;
}

【讨论】:

  • 这可能在这种情况下有效,但通常不能解决问题。也有不适合wchar 的字符。
【解决方案2】:

问题很复杂。非拉丁字符必须正确编码。有几个标准。问题是您的系统正在使用哪种编码。

在 UTF-8 编码中,一个字符由多个字节表示。它可以从 1 到 4 个字节变化,具体取决于它是什么类型的字符。 For example: λ 由两个字节(十六进制)表示:CEBB。

我不知道其他字符编码是什么,它可以为希腊字母提供单字节字符,但我确信有一种这样的编码。

请注意,您的值 my_word.length() 很可能返回 10 而不是 5。

【讨论】:

    【解决方案3】:

    正如其他人所说,这取决于您的编码。一旦您转向国际化,at() 函数就会出现问题,因为例如,希伯来语在字符周围写有元音。并非所有脚本都包含离散的字形序列。

    通常最好将字符串视为原子字符串,除非您自己编写显示/文字操作代码,当然您需要单独的字形。要阅读 UTF,请查看 Baby X 中的代码(它是一个必须在屏幕上绘制文本的窗口系统)

    这里是链接https://github.com/MalcolmMcLean/babyx/blob/master/src/common/BBX_Font.c

    这是 UTF8 代码 - 这是一大堆代码,但基本上很简单。

    static const unsigned int offsetsFromUTF8[6] = 
    {
        0x00000000UL, 0x00003080UL, 0x000E2080UL,
        0x03C82080UL, 0xFA082080UL, 0x82082080UL
    };
    
    static const unsigned char trailingBytesForUTF8[256] = {
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, 0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,
        1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, 1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,
        2,2,2,2,2,2,2,2,2,2,2,2,2,2,2,2, 3,3,3,3,3,3,3,3,4,4,4,4,5,5,5,5
    };
    
    int bbx_isutf8z(const char *str)
    {
      int len = 0;
      int pos = 0;
      int nb;
      int i;
      int ch;
    
      while(str[len])
        len++;
      while(pos < len && *str)
      {
        nb = bbx_utf8_skip(str);
        if(nb < 1 || nb > 4)
          return 0;
        if(pos + nb > len)
          return 0;
        for(i=1;i<nb;i++)
          if( (str[i] & 0xC0) != 0x80 )
            return 0;
        ch = bbx_utf8_getch(str);
        if(ch < 0x80)
        {
          if(nb != 1)
            return 0;
        }
        else if(ch < 0x8000)
        {
          if(nb != 2)
            return 0;
        }
        else if(ch < 0x10000)
        {
          if(nb != 3)
            return 0;
        }
        else if(ch < 0x110000)
        {
          if(nb != 4)
            return 0;
        }
        pos += nb;
        str += nb;    
      }
    
      return 1;
    }
    
    int bbx_utf8_skip(const char *utf8)
    {
      return trailingBytesForUTF8[(unsigned char) *utf8] + 1;
    }
    
    int bbx_utf8_getch(const char *utf8)
    {
        int ch;
        int nb;
    
        nb = trailingBytesForUTF8[(unsigned char)*utf8];
        ch = 0;
        switch (nb) 
        {
                /* these fall through deliberately */
            case 3: ch += (unsigned char)*utf8++; ch <<= 6;
            case 2: ch += (unsigned char)*utf8++; ch <<= 6;
            case 1: ch += (unsigned char)*utf8++; ch <<= 6;
            case 0: ch += (unsigned char)*utf8++;
        }
        ch -= offsetsFromUTF8[nb];
    
        return ch;
    }
    
    int bbx_utf8_putch(char *out, int ch)
    {
      char *dest = out;
      if (ch < 0x80) 
      {
         *dest++ = (char)ch;
      }
      else if (ch < 0x800) 
      {
        *dest++ = (ch>>6) | 0xC0;
        *dest++ = (ch & 0x3F) | 0x80;
      }
      else if (ch < 0x10000) 
      {
         *dest++ = (ch>>12) | 0xE0;
         *dest++ = ((ch>>6) & 0x3F) | 0x80;
         *dest++ = (ch & 0x3F) | 0x80;
      }
      else if (ch < 0x110000) 
      {
         *dest++ = (ch>>18) | 0xF0;
         *dest++ = ((ch>>12) & 0x3F) | 0x80;
         *dest++ = ((ch>>6) & 0x3F) | 0x80;
         *dest++ = (ch & 0x3F) | 0x80;
      }
      else
        return 0;
      return dest - out;
    }
    
    int bbx_utf8_charwidth(int ch)
    {
        if (ch < 0x80)
        {
            return 1;
        }
        else if (ch < 0x800)
        {
            return 2;
        }
        else if (ch < 0x10000)
        {
            return 3;
        }
        else if (ch < 0x110000)
        {
            return 4;
        }
        else
            return 0;
    }
    
    int bbx_utf8_Nchars(const char *utf8)
    {
      int answer = 0;
    
      while(*utf8)
      {
        utf8 += bbx_utf8_skip(utf8);
        answer++;
      }
    
      return answer;
    }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-04-02
      • 2013-09-30
      • 2012-05-16
      • 1970-01-01
      • 2020-04-10
      • 1970-01-01
      相关资源
      最近更新 更多