【问题标题】:Error while converting from std::wstring to wchar_t*从 std::wstring 转换为 wchar_t* 时出错
【发布时间】:2021-06-10 00:23:25
【问题描述】:

我将以下包含 std::wstring 的代码转换为 wchar_t(下一个代码块),但出现内存转储错误:

下面是带有 std::wstring 的代码:

std::wstring RemoveFirstAndLast(std::wstring str)
{
    std::wstring removeFirst = L"";
    bool first = true;
    for (std::wstring::iterator it = str.begin(); it != str.end(); ++it)
    {
        const wchar_t *brckt = L"[";
        const wchar_t *empty = L" ";
        const wchar_t *curlyBrkt = L"{";
        if (((*it) == (*curlyBrkt)) && (first))
        {
            first = false;
            removeFirst = removeFirst + (*it);
        }
        else
        {
            removeFirst = removeFirst + (*it);
        }
    }

    std::wstring removeLast = L"";
    bool last = true;
    for (std::wstring::reverse_iterator it = removeFirst.rbegin(); it != removeFirst.rend(); ++it)
    {
        const wchar_t *brckt = L"]";
        const wchar_t *empty = L" ";
        const wchar_t *curlyBrkt = L"}";

        if (((*it) == (*curlyBrkt)) && (last))
        {
            last = false;
            removeLast = (*it) + removeLast;
        }
        else
        {
            removeLast = (*it) + removeLast;
        }
    }

    return removeLast;
}

下面是与上面相同的代码,但带有 wchar_t*:

wchar_t* RemoveFirstAndLast(wchar_t *str)
{
    wchar_t *removeFirst = L"";
    int length = wcslen(str);
    bool first = true;
    for (int i = 0; i < length; i++)
    {
        const wchar_t *brckt = L"[";
        const wchar_t *empty = L" ";
        const wchar_t *curlyBrkt = L"{";
        if (((str[i]) == (*curlyBrkt)) && (first))
        {
            first = false;
            wcsncat_s(removeFirst, (wcslen(removeFirst) + 1) * 2, &str[i], 1);  //removeFirst = removeFirst + (str[i]);
        }
        else
        {
            wcsncat_s(removeFirst, (wcslen(removeFirst) + 1) * 2, &str[i], 1);  //removeFirst = removeFirst + (str[i]);
        }
    }

    wchar_t *removeLast = L"";
    length = wcslen(removeFirst);
    bool last = true;
    for (int i = length; i >= 0; i--)
    {
        const wchar_t *brckt = L"]";
        const wchar_t *empty = L" ";
        const wchar_t *curlyBrkt = L"}";
        if (((removeFirst[i]) == (*curlyBrkt)) && (last))
        {
            last = false;
            wchar_t *temp2 = (wchar_t*)calloc((wcslen(removeLast) * 2) + 1, sizeof(wchar_t));
            wcsncat_s(temp2, (wcslen(temp2)) * 2, &removeFirst[i], 1);
            wcsncat_s(temp2, (wcslen(temp2)) * 2, removeLast, wcslen(removeLast));
            removeLast = temp2;
        }
        else
        {
            wchar_t *temp3 = (wchar_t*)calloc((wcslen(removeLast) * 2) + 1, sizeof(wchar_t));
            wcsncat_s(temp3, (wcslen(temp3)) * 2, &removeFirst[i], 1);
            wcsncat_s(temp3, (wcslen(temp3)) * 2, removeLast, wcslen(removeLast));
            removeLast = temp3;
            //removeLast = (removeFirst[i]) + removeLast;
        }
    }

    return removeLast;
}

代码正在尝试删除第一个和最后一个字符。作为一个初学者,我试图理解为什么在简单地将代码从 std::wstring 转换为 wchar_t* 时会出现错误?该错误会创建内存转储?有人可以帮我理解我在哪里犯错了吗?

【问题讨论】:

  • 具体是什么“内存转储错误”?
  • 使用您的调试器和编译器的地址清理程序。两者都是有价值的信息来源,也是找出代码存在的任何内存问题的好工具。
  • wcsncat_s(removeFirst (... 从根本上被破坏了,因为您没有为removeFirst 分配任何内存。事实上,我很惊讶 wchar_t *removeFirst = L""; 编译 - 如果将代码编译为 C++,它应该是 const wchar_t *removeFirst = L"";。
  • @PaulSanders:非常感谢……这很有帮助。你能帮我理解如何分配内存,这将非常有用。
  • @chris:这些是保罗所说的错误类型

标签: c++ c++11


【解决方案1】:

正如 Paul 所说 - std::wstring 为您完成所有内存分配和复制/移动。当您使用wchar_t* 时,您正在处理裸指针,这些指针必须寻址您以某种方式分配的内存。请注意,您已经在使用calloc 进行此操作。

在处理宽字符串时,您需要注意计数是“字节数”还是“字符数”,因为它们不相等。最好命名变量以指示它们使用的单位。

赋值wchar_t *removeFirst = L""; 还可以将指针设置为“字符串常量”的地址,该“字符串常量”可能驻留在内存的不可写区域中。在任何情况下,您都无法将任何东西连接到它。相反,您必须提供一个以零宽字符开头的缓冲区,以提供“可以连接的空字符串”。由于您希望从字符串中删除字符,因此您可以分配一个宽度为原始字符串长度的缓冲区,并使用该缓冲区来累积结果。所以你可以做

 size_t original_string_length_widechars = wcslen(str);
 size_t processing_buffer_length_widechars = original_string_length_widechars + 1;
 wchar_t* processing_buf = (wchar_t*)calloc(processing_buffer_length_widechars, sizeof(wchar_t));

添加 1 为零宽字符终止符提供空间(以适应从​​字符串中删除任何内容的情况)。您可能过度分配并不重要(类似于在向量中使用reserve);缓冲区只需要足够大以容纳您正在执行的处理(在这种情况下,从原始字符串复制宽字符,不带第一个和最后一个括号)。

连接函数(例如wcsncat_s)不会为您保留内存;他们只是将字符复制到一个缓冲区,您应该提供足够的空间来容纳连接的结果。

换句话说,您必须对字符串操作有不同的看法。与其考虑“将这两个字符串相加”,在 C 中,您需要考虑“保留一个缓冲区以容纳结果并通过在其中连接字符串来填充它”。

如果您只想复制第一个和最后一个分隔符之间的内容,那么您只需按如下方式填充处理缓冲区:

wchar_t const* position_of_first_char_after_start_delimiter;
wchar_t const* position_of_last_delimiter;

// scan the original str to determine values for these ^^^

// note that pointer arithmetic in C takes the size of the pointer types into account; you don't need to divide this by 2......
size_t num_chars_to_copy = position_of_last_delimiter - position_of_first_char_after_start_delimiter;

wcsncpy_s (processing_buf,
           processing_buffer_length_widechars,
           position_of_first_char_after_start_delimiter,
           num_chars_to_copy);

此外,这部分看起来很狡猾......字符串索引是 0 到 (length-1)。访问 string[length] 可能会引发意外行为。

length = wcslen(removeFirst);
...
for (int i = length; i >= 0; i--)
{
...
 if (((removeFirst[i]) == (*curlyBrkt)) && (last))

此外,在使用宽字符串函数时,您不需要将所有内容“加倍”——另外,在 calloc 中,您正确地声明了每个元素的大小都是宽字符的大小。


函数签名wchar_t* RemoveFirstAndLast(wchar_t *str) 只返回一个裸指针。需要有注释来说明是否为它分配了内存以及调用者是否需要释放它。一般来说,如果可以的话,你应该避免对调用者施加这样的要求(除非它是明确的)。

鉴于该函数旨在修剪已分配的字符串,您可能需要考虑就地修剪str。如果str 无法修改,则需要将其声明为指向常量宽字符的指针,即wchar_t const* str(从右到左读取)。

【讨论】:

  • 假设 length 是字符串长度,那么 string[length] 将是零终止符,访问它应该非常好。
  • 如何分配字符串文字 wchar_t *removeFirst = L"";然后附加到它
  • 你不能像那样“分配一个字符串文字......然后附加到它”,C++ 不能这样工作。
猜你喜欢
  • 2012-02-13
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-09-16
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多