【发布时间】:2017-02-22 16:45:07
【问题描述】:
(我查看了以前的帖子并尝试了他们的建议,但无济于事。)
我正在尝试读取仅包含日文字符的文件。这是该文件的样子:
わたし わ エドワド オ'ハゲンです。
当我尝试读取它时,控制台中没有任何输出显示,而在调试时,读取缓冲区只是垃圾。这是我用来读取文件的函数:
wchar_t* ReadTextFileW(wchar_t* filePath, size_t numBytesToRead, size_t maxBufferSize, const wchar_t* mode, int seekOffset, int seekOrigin)
{
size_t numItems = 0;
size_t bufferSize = 0;
wchar_t* buffer = NULL;
FILE* file = NULL;
//Ensure the filePath does NOT lead to a device.
if (IsPathADevice(filePath) == false)
{
//0 indicates to read as much as possible (the max specified).
if (numBytesToRead == 0)
{
numBytesToRead = maxBufferSize;
}
if (filePath != NULL && mode != NULL)
{
//Ensure there are no errors in opening the file.
if (_wfopen_s(&file, filePath, mode) == 0)
{
//Set the cursor location (back to the beginning of the file by default).
if (fseek(file, seekOffset, seekOrigin) != 0)
{
//Error: Could not change file cursor position.
fclose(file);
return NULL;
}
//Calculate the size of the buffer in bytes.
bufferSize = numBytesToRead * sizeof(wchar_t);
//Create the buffer to store file data in.
buffer = (wchar_t*)_aligned_malloc(bufferSize, BYTE_ALIGNMENT);
//Ensure the buffer was allocated.
if (buffer == NULL)
{
//Error: Buffer could not be allocated.
fclose(file);
return NULL;
}
//Clear any garbage data in the buffer.
memset(buffer, 0, bufferSize);
//Read the data from the file.
numItems = fread_s(buffer, bufferSize, sizeof(wchar_t), numBytesToRead, file);
//Check for read errors.
if (numItems <= 0)
{
//Error: File could not be read.
fclose(file);
_aligned_free(buffer);
return NULL;
}
//Ensure the file is closed without errors.
if (fclose(file) != 0)
{
//Error: File did not close properly.
_aligned_free(buffer);
return NULL;
}
}
}
}
return buffer;
}
要调用此函数,我正在执行以下操作。也许我没有正确使用 setlocale() 但从我读到的内容看来我是。只是重申一下,我遇到的问题是垃圾似乎被读入并且控制台中没有显示任何内容:
setlocale(LC_ALL, "jp");
wchar_t* retVal = ReadTextFileW(L"C:\\jap.txt");
printf("%S\n", retVal);
_aligned_free(retVal);
我还在 .cpp 的顶部定义了以下内容
#define UNICODE
#define _UNICODE
已解决:
如 ryyker 所述,要解决此问题,您需要知道用于创建原始文件的编码。在记事本和记事本++ 中有一个用于编码的下拉菜单。默认情况下(也是最常用的)是 UTF-8。
知道编码后,您可以将_wfopen_s() 的读取模式更改为以下。
wchar_t* retVal = ReadWide::ReadTextFileW(L"C:\\jap.txt", 0, 1024, L"r, ccs=UTF-8");
MessageBoxW(NULL, retVal, NULL, 0);
_aligned_free(retVal);
您必须使用消息框来打印外来字符。
【问题讨论】:
-
您确定文件的encoding 吗?该链接讨论了多字节等。人。
-
@ryyker 我不明白你的问题。
-
我只是建议您尝试解码包含多字节字符的文件,但尚未确定它们由什么编码组成。看看我在上一条评论中留下的链接。
-
你能拿出你的“已解决”部分作为正确答案吗?
-
我看到您找到了解决方案!我刚刚在@Mr.C64 的回答下查看了 cmets 的广泛成绩单。看来 Mr.C64 确实是回答您问题的人。尽管我非常感谢您的 accepted answer 点击,但我认为它应该应用于其他地方。 :)
标签: c windows file file-io unicode