【发布时间】:2017-03-10 23:24:53
【问题描述】:
添加 ru_RU.CP1251 语言环境(在 debian 上取消注释 ru_RU.CP1251 在 /etc/locale.gen 并运行 sudo locale-gen)和
使用gcc -fexec-charset=cp1251 test.c 编译以下程序(输入文件为UTF-8 格式)。结果是空的。只是字母“я”是错误的。
其他字母确定是小写还是大写都可以。
#include <locale.h>
#include <ctype.h>
#include <stdio.h>
int main (void)
{
setlocale(LC_ALL, "ru_RU.CP1251");
char c = 'я';
int i;
char z;
for (i = 7; i >= 0; i--) {
z = 1 << i;
if ((z & c) == z) printf("1"); else printf("0");
}
printf("\n");
if (islower(c))
printf("lowercase\n");
if (isupper(c))
printf("uppercase\n");
return 0;
}
为什么islower() 和isupper() 都不能处理я 的字母?
【问题讨论】:
-
char是否足够大以存储я?islower()的原型表明int会是更好的选择。 -
@KeineLust 我使用 cp1251,它是 8 位编码。我不需要宽字符。试试字母 'ю' - 它工作得很好。只有字母“я”不起作用。我需要解决这个问题。
-
@IgorLiferenko,你试过
ru_RU.UTF-8吗?如果输入是 utf-8,如果您不首先在代码集之间转换字符代码,那么尝试将其显示为 cp-1251 是没有意义的。 -
@mouviciel:常规
isxxxxx()函数的原型具有int作为参数类型,因为您可以传递任何有效的“字符编码为unsigned char”值或 EOF,这意味着一个的char类型不能用作形式参数类型(因为它不能接受足够广泛的值)。