【问题标题】:c++ WriteFile unicode charactersc++ WriteFile unicode 字符
【发布时间】:2015-02-19 23:02:44
【问题描述】:

我正在尝试使用 WriteFile 函数将 wstring 写入 UTF-8 文件。 我希望文件包含这些字符“ÑÁ”,但我得到了这个“�”。

这里是代码

#include <iostream>
#include <cstdlib>
#include <sstream>
#include <string>
#include <fstream>
#include <windows.h>
#include <wchar.h>
#include <stdio.h>
#include <winbase.h>
using namespace std;

const char filepath [] = "unicode.txt";

int main ()
{ 
   wstring str;
   str.append(L"ÑÁ");
   wchar_t* wfilepath;

   // Create a file to work with Unicode and UTF-8
   ofstream fs;
   fs.open(filepath, ios::out|ios::binary);
   unsigned char smarker[3];
   smarker[0] = 0xEF;
   smarker[1] = 0xBB;
   smarker[2] = 0xBF;
   fs << smarker;
   fs.close();

   //Open and write in the file with windows functions
   mbstowcs(wfilepath, filepath, strlen(filepath));
   HANDLE hfile;
   hfile = CreateFileW(TEXT(wfilepath), GENERIC_WRITE, 0, NULL,
       OPEN_EXISTING, FILE_ATTRIBUTE_NORMAL, NULL);
   wstringbuf strBuf (str, ios_base::out|ios::app);
   DWORD bytesWritten;
   DWORD dwBytesToWrite = (DWORD) strBuf.in_avail();
   WriteFile(hfile, &strBuf, dwBytesToWrite, &bytesWritten, NULL);
   CloseHandle(hfile);
}

我在 cygwin 上使用这个命令行编译它:
g++ -std=c++11 -g Windows.C -o Windows

【问题讨论】:

  • 顺便说一句,您不需要使用匈牙利符号,因为 Microsoft 需要。编译器不关心标识符的拼写;只是名称符合 C++ 语言规则。
  • @ThomasMatthews:甚至 MS 都不再使用(反)匈牙利符号了。他们的 .NET 命名指南说不要使用匈牙利符号。
  • 阅读 utf8everywhere.org 了解如何将内容转换为 UTF-8。

标签: c++ c++11 unicode


【解决方案1】:

您需要先将 UTF-16 数据转换为 UTF-8,然后再将其写入文件。

并且无需使用std::ofstream 创建文件,关闭它,然后使用CreateFileW() 重新打开它。只需打开文件一次,然后写下您需要的所有内容。

试试这个:

#include <iostream>
#include <cstdlib>
#include <string>
//#include <codecvt>
//#include <locale>

#include <windows.h>
#include <wchar.h>
#include <stdio.h>

using namespace std;

LPCWSTR filepath = L"unicode.txt";

string to_utf8(const wstring &s)
{
    /*
    wstring_convert<codecvt_utf8_utf16<wchar_t>> utf16conv;
    return utf16conv.to_bytes(s);
    */

    string utf8;
    int len = WideCharToMultiByte(CP_UTF8, 0, s.c_str(), s.length(), NULL, 0, NULL, NULL);
    if (len > 0)
    {
        utf8.resize(len);
        WideCharToMultiByte(CP_UTF8, 0, s.c_str(), s.length(), &utf8[0], len, NULL, NULL);
    }
    return utf8;
}

int main ()
{ 
    wstring str = L"ÑÁ";

    // Create a UTF-8 file and write in it using Windows functions
    HANDLE hfile = CreateFileW(filepath, GENERIC_WRITE, 0, NULL,
       CREATE_ALWAYS, FILE_ATTRIBUTE_NORMAL, NULL);
    if (hfile != INVALID_HANDLE_VALUE)
    {
        unsigned char smarker[3];
        DWORD bytesWritten;

        smarker[0] = 0xEF;
        smarker[1] = 0xBB;
        smarker[2] = 0xBF;
        WriteFile(hfile, smarker, 3, &bytesWritten, NULL);

        string strBuf = to_utf8(str);
        WriteFile(hfile, strBuf.c_str(), strBuf.size(), &bytesWritten, NULL);

        CloseHandle(hfile);
    }

    return 0;
}

【讨论】:

    【解决方案2】:

    问题就在这里:

    wstringbuf strBuf (str, ios_base::out|ios::app);
    WriteFile(hfile, &strBuf, dwBytesToWrite, &bytesWritten, NULL);
    

    &amp;strBuf 是wstringbuf 对象的地址,其中包含指向内容的指针、缓冲区位置和状态标志……而不是其内容所在的位置。

    你可能想要

    WriteFile(hfile, &str[0], /* etc */
    

    但这只会存储您的wstring 使用的相同编码。要使用 UTF-8 编写,您可能需要使用 WideCharToMultiByte(或 wcstombs,因为您已经使用了 mbstowcs)。

    【讨论】:

    • 嗨,我试过这个char* strBuf; wcstombs(strBuf, str.c_str(), wcslen(str.c_str())); WriteFile(hfile, &amp;strBuf[0],...) 但是输出文件是空白的我做错了什么
    • wcstombs 不分配内存,你需要给它一个指向足够大缓冲区的指针。
    • 重要的一课,但是我不明白为什么 Windows 想让你从 wchar 更改为 multibyte...
    • @NorbertBoros:因为 UTF-8 是可变长度编码?
    【解决方案3】:

    Ben 是正确的,您将原始 wchar_t 写入文件,而不是 UTF-8。

    要编写 UTF-8,您可以考虑留在 C++ 中并这样做:

    std::locale loc (std::locale(), new std::codecvt_utf8<wchar_t>);
    
    std::wofstream fs ("unicode.txt");
    fs.imbue(loc);
    
    fs << L"ÑÁ"; 
    

    【讨论】:

    猜你喜欢
    • 2013-11-02
    • 1970-01-01
    • 2013-06-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-12-06
    相关资源
    最近更新 更多