【问题标题】:C: Select a substring, n columns wide from a multibyte stringC:从多字节字符串中选择一个子字符串,n 列宽
【发布时间】:2014-10-24 02:04:43
【问题描述】:

我在 C 中有一个 char * 字符串,它基于用户输入。从这个字符串中,我想从第一个位置开始选择一个子字符串,这样生成的子字符串在固定宽度的终端上是 n 列宽。

过去从未使用过非 ASCII 字符,我完全不知道如何解决这个问题,甚至开始。一些初步搜索建议使用libiconv,但这似乎没有帮助。我也尝试使用wchar.h,广泛的字符支持,但我不确定这是正确的方法。

编辑:这是我第一次尝试的尝试:

static int
count_n_cols (const char *mbs, char *mbf, const int n)
{
  wchar_t wc;
  int     bytes;
  int     remaining = strlen(mbs);
  int     cols = 0;
  int     wccols;

  while (*mbs != '\0' && cols <= n)
    {
      bytes = mbtowc (&wc, mbs, remaining);
      assert (bytes != 0);  /* Only happens when *mbs == '\0' */
      if (bytes == -1)
        {
          /* Invalid sequence. We'll just have to fudge it. */
          return cols + remaining;
        }
      mbs += bytes;
      remaining -= bytes;
      wccols = wcwidth(wc);
      *mbf += wc;
      cols += (wccols == -1? 1 : wccols);
    }
  return cols;
}

【问题讨论】:

  • 添加了我尝试编写的函数。我不确定这是否是最好的前进方式。
  • 您的代码有什么问题?乍一看还不错。

标签: c string multibyte


【解决方案1】:

如果我正确理解您的问题,您想计算 utf-8 序列的数量以执行您的子字符串而不进行任何转换。您可以按照 utf-8 标准的规定,通过读取序列的第一个字节来计算每个“列”对应的字节数。这是一些示例代码,基于您的示例函数和Wikipedia's UTF-8 description:

static int count_n_cols (const char *mbs, char *mbf, const int n)
{
    int bytes;
    int length = strlen(mbs);
    int cols = 0;

    for (bytes = 0; bytes < length; bytes++)
    {
        if (mbs[bytes] == '\0' || cols >= n)
            break;
        else if ((mbs[bytes] & 0x80) == 0)  // the first bit is 0
        {
            cols++;
        }
        else if ((mbs[bytes] & 0xE0) == 0xC0)   //the first 3 bits are 110
        {
            //two bytes in utf8 sequence
            cols++;
            bytes++;
        }
        else if ((mbs[bytes] & 0xF0) == 0xE0)   //the first 4 bits are 1110
        {
            //three bytes in utf8 sequence
            cols++;
            bytes += 2;
        else if ((mbs[bytes] & 0xF8) == 0xF0)   //the first 5 bits are 11110
        {
            //four bytes in utf8 sequence
            cols++;
            bytes += 3;
        }
        else
        {
            putc(mbs[bytes],stdout);
            printf(" non_ascii %d\n", mbs[bytes] & 0x80);
        }
    }
    strncpy(mbf, mbs, bytes);
    mbf[bytes] = '\0';
    return cols;
}

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-07-12
    • 2018-05-25
    • 1970-01-01
    • 1970-01-01
    • 2021-05-26
    • 1970-01-01
    相关资源
    最近更新 更多