【问题标题】:What is exact analog of C#'s string.Compare ignoring case in C++?C# 的 string.Compare 忽略 C++ 中的大小写的确切模拟是什么?
【发布时间】:2022-01-10 23:49:53
【问题描述】:

也许有人知道,C# 的 string.Compare 忽略大小写的确切 C 或 C++(任何一个都可以)模拟是什么? 事实证明,_wcsicmp 不同,尽管两者都应该使用当前的语言环境或文化(即en_US)。

With string.Compare(..., true),
 or string.Compare(..., StringComparison.CurrentCultureIgnoreCase),
 or string.Compare(..., StringComparison.InvariantCultureIgnoreCase):
'~' before '+',
'=' before number,
letter before single quote

_wcsicmpwcsicmp_l 具有显式语言环境 (LC_ALL, L"en_US") 将它们按相反的顺序排列。与std::wcscoll 的结果完全相同。

我可以使用字符表重现它,但也许有更好的方法。 谢谢!

===== 可能没人知道。我正在发布解决方法,这对于 C# 是不必要的。它负责 ANSI 子集(0-256,我最关心)和部分 Unicode 表的其余部分:

int compareNoCase(const std::wstring& a, const std::wstring& b, int size = -1)
{
    return compareNoCase(a.c_str(), b.c_str(), size);
}

int compareNoCase(LPCWSTR a, LPCWSTR b, int size = -1)
{
    static const unsigned char table[] = {
        0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08, 0x09, 0x0a, 0x0b, 0x0c, 0x0d, 0x0e, 0x0f,
        0x10, 0x11, 0x12, 0x13, 0x14, 0x15, 0x16, 0x17, 0x18, 0x19, 0x1a, 0x1b, 0x1c, 0x1d, 0x1e, 0x1f,
        0x20, 0x21, 0x22, 0x23, 0x24, 0x25, 0x26, 0x63, 0x27, 0x28, 0x29, 0x3d, 0x2a, 0x64, 0x2b, 0x2c,
        0x3f, 0x40, 0x41, 0x42, 0x43, 0x44, 0x45, 0x46, 0x47, 0x48, 0x2d, 0x2e, 0x2f, 0x3e, 0x30, 0x31,
        0x32, 0x49, 0x4a, 0x4b, 0x4c, 0x4d, 0x4e, 0x4f, 0x50, 0x51, 0x52, 0x53, 0x54, 0x55, 0x56, 0x57,
        0x58, 0x59, 0x5a, 0x5b, 0x5c, 0x5d, 0x5e, 0x5f, 0x60, 0x61, 0x62, 0x33, 0x34, 0x35, 0x36, 0x37,
        0x38, 0x49, 0x4a, 0x4b, 0x4c, 0x4d, 0x4e, 0x4f, 0x50, 0x51, 0x52, 0x53, 0x54, 0x55, 0x56, 0x57,
        0x58, 0x59, 0x5a, 0x5b, 0x5c, 0x5d, 0x5e, 0x5f, 0x60, 0x61, 0x62, 0x39, 0x3a, 0x3b, 0x3c, 0x65
    };

    for (int i = 0; size < 0 || i < size; i++) {
        wchar_t ca = a[i];
        wchar_t cb = b[i];
        if (ca == 0 || cb == 0) {           // if at least one of the strings is over:
            return (ca == 0) ? ((cb == 0) ? 0 : -1) : 1;
        }
        if (ca != cb) {                     // if next characters are different, go in
            if (ca < 0x7f && cb < 0x7f) {   // if both characters are ASCII, use table
                if (table[ca] != table[cb]) {
                    return (table[ca] > table[cb]) ? 1 : -1;
                }
            }
            else {                          // otherwise use default system locale
                int ret = std::wcscoll(a + i, b + i);
                if (ret != 0) {
                    return ret;
                }
            }
        }
    }
    return 0;
}

表格不包含字符,文件名中禁止使用。 Explorer 风格的“数字”比较与这个问题无关。为了清楚起见,我还删除了对多个语言环境的处理。

如果有人有更好的想法,请告诉我!

【问题讨论】:

  • 没有像 C/C++ 这样的语言。它们是两种不同的语言,两者的答案会有所不同。请只选择一个。
  • @kaylum - 好的,C 或 C++。我知道它们是不同的语言,但我很乐意在其中任何一个中得到答案。否则我会更具体。
  • @kaylum,你能回答 C 语言的问题吗?对于 C++ 语言?
  • 有点下结论?我没有投反对票。而且我不需要能够回答问题来提出改进问题以符合 Stack Overflow 准则的方法。
  • @kaylum,我很抱歉。问题是,如果问题很简单并且完全可以查找,那么它就会得到极大的支持并且很容易回答。混乱的术语或糟糕的英语是可以原谅的。而且,相反,如果这个问题似乎遥不可及,它会立即被否决。出于某种原因,我的“C”标签被删除了(可能是用“C++”折叠起来的)。我仅在 C 中提供了示例,因为它似乎很容易显示我正在寻找的内容。

标签: c# c++ c string compare


【解决方案1】:

在 C 中,strings.h 中有一个用于此的函数,称为 strcasecmp()。尽管您可能需要一个等效的(VS 中的 Windows),如此答案中所述:error C3861: 'strcasecmp': identifier not found in visual studio 2008?

所以你可以写这样的东西。

#include <stdio.h>
#include <strings.h>

#ifdef _MSC_VER
//not #if defined(_WIN32) || defined(_WIN64) because we have strncasecmp in mingw
#define strncasecmp _strnicmp
#define strcasecmp _stricmp
#endif

int main () {
    char *str1 = "this is a string";
    char *str2 = "THIS IS A STRING";

    if (strcasecmp(str1, str2) == 0) printf("Strings match\n");

    return 0;
}

如果你真的想要,你可以通过使用它们的c_str() 函数来对 C++ 字符串使用相同的方法。

strcasecmp(str1.c_str(), str2.c_str())

但实际上你最好使用boost::iequals(str1, str2)

#include <stdio.h>
#include <boost/algorithm/string.hpp>

int main () {
    std::string str1 = "this is a string";
    std::string str2 = "THIS IS A STRING";

    if (boost::iequals(str1, str2)) printf("Strings match\n");

    return 0;
}

【讨论】:

  • 谢谢!它归结为相同的 _stricmp,它是区域感知的,这与 C# 使用的有所不同。
  • 我认为语言环境至少在 Windows 环境中被用作“黄金标准”,但看起来 .NET 做了一些不完全兼容的事情。区别在于 3 种情况,涉及特殊字符。顺便说一句,C# 的排序对应于文件资源管理器的文件名排序(如果忽略连续数字,则另当别论)。我很好奇,是否有一种方法可以在不借助字符表的情况下通过 C 或 C++ 合法地实现排序顺序。
  • C# 也有区域感知比较。它只是使用您的默认文化。您真正需要的是获得正确的语言环境。您需要阅读有关 C# 中的文化和本地人以及 Windows API 中的 C 或 C++ 的文档。
  • @siride,我尝试使用 将其作为小写进行比较,但它不起作用。也许事情是与 std::locale 一起玩。理论上它应该使用默认语言环境(en_US)。规范没有说明 _locale_t 在功能上与 std::locale 相同。会检查的,谢谢!
  • 我查看了 String 的源代码,最终它调用了一些 C++ 库函数:referencesource.microsoft.com/#mscorlib/system/globalization/…。这是:stackoverflow.com/a/23118596/394487。所以也许你可以使用那个 COM 函数。不过,我似乎找不到任何文档。
猜你喜欢
  • 1970-01-01
  • 2010-12-07
  • 1970-01-01
  • 2010-11-26
  • 1970-01-01
  • 2018-02-10
  • 1970-01-01
相关资源
最近更新 更多