【问题标题】:Encoding issue with spanish file in C#C#中西班牙语文件的编码问题
【发布时间】:2021-04-14 02:13:32
【问题描述】:

我在西班牙语的 azure blob 存储中有一个在线文件存储。有些单词有特殊字符(例如:Almacén) 当我在 notepad++ 中打开文件时,编码是 ANSI。

所以现在我尝试使用代码读取文件:

        using StreamReader reader = new StreamReader(Stream, Encoding.UTF8);
        blobStream.Seek(0, SeekOrigin.Begin);
        var allLines = await reader.ReadToEndAsync();

问题是“allLines”编码不正确,我有一些问题,例如:Almac�n

我已经尝试了一些像这样的解决方案: C# Convert string from UTF-8 to ISO-8859-1 (Latin1) H

还是不行

(最终目标是“合并”两个csv,所以我读取了两者的流,删除标题并连接字符串以再次推送它。如果有更好的解决方案可以在c#中合并csv,可以跳过此编码问题我也愿意)

【问题讨论】:

    标签: c# string encoding


    【解决方案1】:

    您正在尝试读取非 UTF8 编码的文件,就好像它是 UTF8 编码的一样。我可以用

    复制这个问题
    var s = "Almacén";
    using var memStream = new MemoryStream(Encoding.GetEncoding(28591).GetBytes(s));
    
    using var reader = new StreamReader(memStream, Encoding.UTF8);
    var allLines = await reader.ReadToEndAsync();
    
    Console.WriteLine(allLines); // writes "Almac�n" to console
    

    您应该尝试读取编码为 iso-8859-1 "Western European (ISO)" 的文件,即代码页 28591。

    using var reader = new StreamReader(Stream, Encoding.GetEncoding(28591));
    var allLines = await reader.ReadToEndAsync();
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-12-29
      • 2015-08-07
      • 2013-07-22
      • 2017-08-22
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多