【问题标题】:Writing utf16 to file in binary mode以二进制模式将 utf16 写入文件
【发布时间】:2008-10-16 07:17:39
【问题描述】:

我正在尝试使用 ofstream 以二进制模式将 wstring 写入文件,但我认为我做错了什么。这是我尝试过的:

ofstream outFile("test.txt", std::ios::out | std::ios::binary);
wstring hello = L"hello";
outFile.write((char *) hello.c_str(), hello.length() * sizeof(wchar_t));
outFile.close();

在例如 Firefox 中打开 test.txt,编码设置为 UTF16,它将显示为:

嘿嘿嘿

谁能告诉我为什么会这样?

编辑:

在十六进制编辑器中打开文件我得到:

FF FE 68 00 00 00 65 00 00 00 6C 00 00 00 6C 00 00 00 6F 00 00 00 

由于某种原因,我似乎在每个字符之间多了两个字节?

【问题讨论】:

  • 向与流关联的本地添加一个方面,以执行从 wchar_t 到正确输出的转换。见下文。

标签: c++ unicode utf-16


【解决方案1】:

在这里,我们遇到了很少使用的语言环境属性。 如果您将字符串输出为字符串(而不是原始数据),则可以让语言环境自动进行适当的转换。

注意此代码未考虑 wchar_t 字符的 edianness。

#include <locale>
#include <fstream>
#include <iostream>
// See Below for the facet
#include "UTF16Facet.h"

int main(int argc,char* argv[])
{
   // construct a custom unicode facet and add it to a local.
   UTF16Facet *unicodeFacet = new UTF16Facet();
   const std::locale unicodeLocale(std::cout.getloc(), unicodeFacet);

   // Create a stream and imbue it with the facet
   std::wofstream   saveFile;
   saveFile.imbue(unicodeLocale);


   // Now the stream is imbued we can open it.
   // NB If you open the file stream first. Any attempt to imbue it with a local will silently fail.
   saveFile.open("output.uni");
   saveFile << L"This is my Data\n";


   return(0);
}    

文件:UTF16Facet.h

 #include <locale>

class UTF16Facet: public std::codecvt<wchar_t,char,std::char_traits<wchar_t>::state_type>
{
   typedef std::codecvt<wchar_t,char,std::char_traits<wchar_t>::state_type> MyType;
   typedef MyType::state_type          state_type;
   typedef MyType::result              result;


   /* This function deals with converting data from the input stream into the internal stream.*/
   /*
    * from, from_end:  Points to the beginning and end of the input that we are converting 'from'.
    * to,   to_limit:  Points to where we are writing the conversion 'to'
    * from_next:       When the function exits this should have been updated to point at the next location
    *                  to read from. (ie the first unconverted input character)
    * to_next:         When the function exits this should have been updated to point at the next location
    *                  to write to.
    *
    * status:          This indicates the status of the conversion.
    *                  possible values are:
    *                  error:      An error occurred the bad file bit will be set.
    *                  ok:         Everything went to plan
    *                  partial:    Not enough input data was supplied to complete any conversion.
    *                  nonconv:    no conversion was done.
    */
   virtual result  do_in(state_type &s,
                           const char  *from,const char *from_end,const char* &from_next,
                           wchar_t     *to,  wchar_t    *to_limit,wchar_t*    &to_next) const
   {
       // Loop over both the input and output array/
       for(;(from < from_end) && (to < to_limit);from += 2,++to)
       {
           /*Input the Data*/
           /* As the input 16 bits may not fill the wchar_t object
            * Initialise it so that zero out all its bit's. This
            * is important on systems with 32bit wchar_t objects.
            */
           (*to)                               = L'\0';

           /* Next read the data from the input stream into
            * wchar_t object. Remember that we need to copy
            * into the bottom 16 bits no matter what size the
            * the wchar_t object is.
            */
           reinterpret_cast<char*>(to)[0]  = from[0];
           reinterpret_cast<char*>(to)[1]  = from[1];
       }
       from_next   = from;
       to_next     = to;

       return((from > from_end)?partial:ok);
   }



   /* This function deals with converting data from the internal stream to a C/C++ file stream.*/
   /*
    * from, from_end:  Points to the beginning and end of the input that we are converting 'from'.
    * to,   to_limit:  Points to where we are writing the conversion 'to'
    * from_next:       When the function exits this should have been updated to point at the next location
    *                  to read from. (ie the first unconverted input character)
    * to_next:         When the function exits this should have been updated to point at the next location
    *                  to write to.
    *
    * status:          This indicates the status of the conversion.
    *                  possible values are:
    *                  error:      An error occurred the bad file bit will be set.
    *                  ok:         Everything went to plan
    *                  partial:    Not enough input data was supplied to complete any conversion.
    *                  nonconv:    no conversion was done.
    */
   virtual result do_out(state_type &state,
                           const wchar_t *from, const wchar_t *from_end, const wchar_t* &from_next,
                           char          *to,   char          *to_limit, char*          &to_next) const
   {
       for(;(from < from_end) && (to < to_limit);++from,to += 2)
       {
           /* Output the Data */
           /* NB I am assuming the characters are encoded as UTF-16.
            * This means they are 16 bits inside a wchar_t object.
            * As the size of wchar_t varies between platforms I need
            * to take this into consideration and only take the bottom
            * 16 bits of each wchar_t object.
            */
           to[0]     = reinterpret_cast<const char*>(from)[0];
           to[1]     = reinterpret_cast<const char*>(from)[1];

       }
       from_next   = from;
       to_next     = to;

       return((to > to_limit)?partial:ok);
   }
};

【讨论】:

  • 请注意,您的 Facet 实现与 UCS-2 的转换,而不是 UTF-16。 UTF-16 是一种可变长度编码,用于检测所谓的代理对。 UCS-2 是 Unicode 的一个子集,这就是发明 UTF-16 的原因。
【解决方案2】:

我怀疑 sizeof(wchar_t) 在您的环境中为 4 - 即它写出 UTF-32/UCS-4 而不是 UTF-16。这肯定是十六进制转储的样子。

这很容易测试(只需打印出 sizeof(wchar_t)),但我很确定这是怎么回事。

要从 UTF-32 wstring 转换为 UTF-16,您需要应用适当的编码,因为代理对开始发挥作用。

【讨论】:

  • 是的,你是正确的 wchar_t 大小为 4,我在 mac 中。所以这解释了很多 :) 我知道 UTF-16 中的代理对,将不得不进一步研究。
  • 从输出中你无法判断它是 UTF-16 还是 UTF-32,它只显示 wchar_t 是 4 字节宽。字符串的编码不是由语言定义的(尽管它很可能是 UCS-4)。
【解决方案3】:

如果您使用C++11 标准,这很容易(因为有很多额外的包含,例如"utf8",可以永远解决这个问题)。

但是如果你想使用旧标准的多平台代码,你可以使用这种方法来编写流:

  1. Read the article about UTF converter for streams
  2. 从以上来源将stxutif.h 添加到您的项目中
  3. 以 ANSI 模式打开文件并将 BOM 添加到文件的开头,如下所示:

    std::ofstream fs;
    fs.open(filepath, std::ios::out|std::ios::binary);
    
    unsigned char smarker[3];
    smarker[0] = 0xEF;
    smarker[1] = 0xBB;
    smarker[2] = 0xBF;
    
    fs << smarker;
    fs.close();
    
  4. 然后以UTF 的身份打开文件并在其中写入您的内容:

    std::wofstream fs;
    fs.open(filepath, std::ios::out|std::ios::app);
    
    std::locale utf8_locale(std::locale(), new utf8cvt<false>);
    fs.imbue(utf8_locale); 
    
    fs << .. // Write anything you want...
    

【讨论】:

    【解决方案4】:

    在使用 wofstream 和上面定义的 utf16 方面的 Windows 上失败,因为 wofstream 将所有值为 0A 的字节转换为 2 个字节 0D 0A,这与您如何传递 0A 字节,'\x0A',L'\ x0A'、L'\x000A'、'\n'、L'\n' 和 std::endl 都给出相同的结果。 在 Windows 上,您必须以二进制模式使用 ofstream(不是 wofsteam)打开文件,并像在原始帖子中一样写入输出。

    【讨论】:

      【解决方案5】:

      提供的Utf16Facetgcc 中不适用于大字符串,这是适合我的版本...这样文件将保存在UTF-16LE 中。对于UTF-16BE,只需反转do_indo_out 中的分配,例如to[0] = from[1]to[1] = from[0]

      #include <locale>
      #include <bits/codecvt.h>
      
      
      class UTF16Facet: public std::codecvt<wchar_t,char,std::char_traits<wchar_t>::state_type>
      {
         typedef std::codecvt<wchar_t,char,std::char_traits<wchar_t>::state_type> MyType;
         typedef MyType::state_type          state_type;
         typedef MyType::result              result;
      
      
         /* This function deals with converting data from the input stream into the internal stream.*/
         /*
          * from, from_end:  Points to the beginning and end of the input that we are converting 'from'.
          * to,   to_limit:  Points to where we are writing the conversion 'to'
          * from_next:       When the function exits this should have been updated to point at the next location
          *                  to read from. (ie the first unconverted input character)
          * to_next:         When the function exits this should have been updated to point at the next location
          *                  to write to.
          *
          * status:          This indicates the status of the conversion.
          *                  possible values are:
          *                  error:      An error occurred the bad file bit will be set.
          *                  ok:         Everything went to plan
          *                  partial:    Not enough input data was supplied to complete any conversion.
          *                  nonconv:    no conversion was done.
          */
         virtual result  do_in(state_type &s,
                                 const char  *from,const char *from_end,const char* &from_next,
                                 wchar_t     *to,  wchar_t    *to_limit,wchar_t*    &to_next) const
         {
      
             for(;from < from_end;from += 2,++to)
             {
                 if(to<=to_limit){
                     (*to)                               = L'\0';
      
                     reinterpret_cast<char*>(to)[0]  = from[0];
                     reinterpret_cast<char*>(to)[1]  = from[1];
      
                     from_next   = from;
                     to_next     = to;
                 }
             }
      
             return((to != to_limit)?partial:ok);
         }
      
      
      
         /* This function deals with converting data from the internal stream to a C/C++ file stream.*/
         /*
          * from, from_end:  Points to the beginning and end of the input that we are converting 'from'.
          * to,   to_limit:  Points to where we are writing the conversion 'to'
          * from_next:       When the function exits this should have been updated to point at the next location
          *                  to read from. (ie the first unconverted input character)
          * to_next:         When the function exits this should have been updated to point at the next location
          *                  to write to.
          *
          * status:          This indicates the status of the conversion.
          *                  possible values are:
          *                  error:      An error occurred the bad file bit will be set.
          *                  ok:         Everything went to plan
          *                  partial:    Not enough input data was supplied to complete any conversion.
          *                  nonconv:    no conversion was done.
          */
         virtual result do_out(state_type &state,
                                 const wchar_t *from, const wchar_t *from_end, const wchar_t* &from_next,
                                 char          *to,   char          *to_limit, char*          &to_next) const
         {
      
             for(;(from < from_end);++from, to += 2)
             {
                 if(to <= to_limit){
      
                     to[0]     = reinterpret_cast<const char*>(from)[0];
                     to[1]     = reinterpret_cast<const char*>(from)[1];
      
                     from_next   = from;
                     to_next     = to;
                 }
             }
      
             return((to != to_limit)?partial:ok);
         }
      };
      

      【讨论】:

        【解决方案6】:

        您应该在十六进制编辑器(例如WinHex)中查看输出文件,以便查看实际的位和字节,以验证输出实际上是 UTF-16。把它贴在这里,让我们知道结果。这将告诉我们是责怪 Firefox 还是你的 C++ 程序。

        但在我看来,您的 C++ 程序可以正常工作,而 Firefox 没有正确解释您的 UTF-16。 UTF-16 要求每个字符占用两个字节。但是 Firefox 打印的字符数是它应该打印的两倍,因此它可能试图将您的字符串解释为 UTF-8 或 ASCII,通常每个字符只有 1 个字节。

        当您说“Firefox 编码设置为 UTF16”时,您是什么意思?我怀疑这项工作是否有效。

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2017-09-21
          • 1970-01-01
          • 1970-01-01
          • 2011-04-14
          • 1970-01-01
          • 1970-01-01
          • 2013-07-10
          相关资源
          最近更新 更多