【问题标题】:Converting ANSI to UTF-8 & java.lang.OutOfMemoryError: Java heap space将 ANSI 转换为 UTF-8 和 java.lang.OutOfMemoryError:Java 堆空间
【发布时间】:2018-10-24 09:06:30
【问题描述】:

我的最终目标是将文件从 ANSI 转换为 UTF-8。为此,我使用了一些 Java 代码:

import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.CharBuffer;
import java.nio.charset.Charset;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;

public class ConvertFromAnsiToUtf8 {

    public static void main(String[] args) throws IOException {

        try {
            Path p = Paths.get("C:\\shared_to_vm\\test_encode\\test.csv");
            ByteBuffer bb = ByteBuffer.wrap(Files.readAllBytes(p));
            CharBuffer cb = Charset.forName("windows-1252").decode(bb);
            bb = Charset.forName("UTF-8").encode(cb);
            Files.write(p, bb.array());
        } catch (Exception e) {
            System.out.println(e);
        } 

    } 

}

当我在小文件上测试代码时,它可以完美运行。我的文件从 ANSI 转换为 UTF-8,所有字符都被识别和编码。但是,一旦我尝试在需要转换的文件上使用它,就会收到错误 java.lang.OutOfMemoryError: Java heap space。

就我的理解而言,我的文件中有大约 150 万行,所以我很确定我用我的应用程序创建了太多的对象。

当然,我已经检查过这个错误的含义以及如何解决它(例如here 或here),但是提高我的 JVM 的内存容量是解决它的唯一方法吗?如果是,我应该使用多少?

任何形式的帮助(提示、建议、链接或其他)将不胜感激!

【问题讨论】:

    标签: java utf-8 heap-memory ansi


    【解决方案1】:

    不要一次读取整个文件:

    ByteBuffer bb = ByteBuffer.wrap(Files.readAllBytes(p));
    

    请尝试逐行阅读:

    Files.lines(p, Charset.forName("windows-1252")).forEach(line -> {
       // Convert your line, write to file
    });
    

    【讨论】:

    • 这无法关闭输入流。您必须根据Files.lines() 的结果调用Stream.close()。这种方法是“有损的”,因为它可以改变行尾。
    【解决方案2】:

    如果您有一个比可用随机存取内存大的大文件,您应该逐块转换字符。

    您可以找到以下示例:

    import java.io.IOException;
    import java.nio.ByteBuffer;
    import java.nio.channels.FileChannel;
    import java.nio.channels.ReadableByteChannel;
    import java.nio.channels.WritableByteChannel;
    import java.nio.charset.Charset;
    import java.nio.charset.CharsetDecoder;
    import java.nio.charset.CharsetEncoder;
    import java.nio.file.Path;
    import java.nio.file.Paths;
    import java.nio.file.StandardOpenOption;
    
    public class Iconv {
    
        private static void iconv(Charset toCode, Charset fromCode, Path src, Path dst) throws IOException {
            CharsetDecoder decoder = fromCode.newDecoder();
            CharsetEncoder encoder = toCode.newEncoder();
            try (ReadableByteChannel source = FileChannel.open(src, StandardOpenOption.READ);
                    WritableByteChannel destination = FileChannel.open(dst, StandardOpenOption.CREATE, StandardOpenOption.TRUNCATE_EXISTING,
                            StandardOpenOption.WRITE);) {
                ByteBuffer readBytes = ByteBuffer.allocate(4096);
                while (source.read(readBytes) > 0) {
                    readBytes.flip();
                    destination.write(encoder.encode(decoder.decode(readBytes)));
                    readBytes.clear();
                }
            }
        }
    
        public static void main(String[] args) throws Exception {
            iconv(Charset.forName("UTF-8"), Charset.forName("Windows-1252"), Paths.get("test.csv") , Paths.get("test-utf8.csv") );
        }
    
    }
    

    【讨论】:

    • 谢谢维克多。我尝试了您的解决方案,因为它对我来说看起来更容易,而且似乎效果很好!你在这里真的帮了我!
    【解决方案3】:

    流式传输输入,转换字符编码,并随时写入输出。这样,您不需要将整个文件读入内存,而只需要尽可能多地读入内存即可。

    如果您想最大限度地减少(缓慢)系统调用的数量,您可以使用类似的方法,但显式创建具有更大内部缓冲区的BufferedInputStream,然后将其包装在InputStreamReader 中。但是这里显示的简单方法在许多应用程序中不太可能成为关键点。

    private static final Charset WINDOWS1252 = Charset.forName("windows-1252");
    
    private static final int DEFAULT_BUF_SIZE = 8192;
    
    public static void transcode(Path input, Path output) throws IOException {
        try (Reader r = Files.newBufferedReader(input, WINDOWS1252);
             Writer w = Files.newBufferedWriter(output, StandardCharsets.UTF_8, StandardOpenOption.CREATE_NEW)) {
            char[] buf = new char[DEFAULT_BUF_SIZE];
            while (true) {
                int n = r.read(buf);
                if (n < 0) break;
                w.write(buf, 0, n);
            }
        }
    }
    

    【讨论】:

    • 感谢您的帮助,埃里克森。您的意见确实帮助我掌握了我的问题。我唯一没有完全明白的是在哪里为我的输入和输出编写路径。我试图这样做: Path input = Paths.get("/myInputPath/myAnsiFile.csv");路径输出 = Paths.get("/myOutputPath/myUtf8File.csv");但它没有用。对不起,如果我的问题听起来很无聊。我还是 Java 新手...
    • @DavidJoshua 也许是因为您尝试写入的文件已经存在。为了安全起见,我添加了一个选项 (StandardOpenOption.CREATE_NEW),它可以防止新文件覆盖现有文件。
    猜你喜欢
    • 1970-01-01
    • 2019-02-10
    • 1970-01-01
    • 2015-10-06
    • 2015-07-07
    • 2013-12-14
    • 2012-01-08
    • 2015-11-22
    • 2013-02-11
    相关资源
    最近更新 更多