【问题标题】:Read UTF8/UNICODE characters from an escaped ASCII sequence从转义的 ASCII 序列中读取 UTF8/UNICODE 字符
【发布时间】:2012-12-07 13:35:16
【问题描述】:

我在一个文件中有以下名称,我需要将该字符串读取为 UTF8 编码的字符串,所以从这里开始:

test_\303\246\303\270\303\245.txt

我需要获取以下内容:

test_æøå.txt

你知道如何使用 C# 实现这一点吗?

【问题讨论】:

    标签: c# .net unicode encoding utf-8


    【解决方案1】:

    假设你有这个字符串:

    string input = "test_\\303\\246\\303\\270\\303\\245.txt";
    

    I.E.从字面上看

    test_\303\246\303\270\303\245.txt
    

    你可以这样做:

    string input = "test_\\303\\246\\303\\270\\303\\245.txt";
    Encoding iso88591 = Encoding.GetEncoding(28591); //See note at the end of answer
    Encoding utf8 = Encoding.UTF8;
    
    
    //Turn the octal escape sequences into characters having codepoints 0-255
    //this results in a "binary string"
    string binaryString = Regex.Replace(input, @"\\(?<num>[0-7]{3})", delegate(Match m)
    {
        String oct = m.Groups["num"].ToString();
        return Char.ConvertFromUtf32(Convert.ToInt32(oct, 8));
    
    });
    
    //Turn the "binary string" into bytes
    byte[] raw = iso88591.GetBytes(binaryString);
    
    //Read the bytes into C# string
    string output = utf8.GetString(raw);
    Console.WriteLine(output);
    //test_æøå.txt
    

    “二进制字符串”是指仅由代码点为 0-255 的字符组成的字符串。因此,它相当于一个穷人的byte[],其中 您在索引i 处检索字符的代码点,而不是在索引i 处的byte[] 中的byte 值(这是我们几年前在javascript 中所做的)。因为 iso-8859-1 映射 正是前 256 个 unicode 代码点转换为一个字节,非常适合将“二进制字符串”转换为 byte[]

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2011-02-19
      • 1970-01-01
      • 2010-12-09
      • 2017-01-16
      • 2013-07-19
      相关资源
      最近更新 更多