【问题标题】:Java NIO scan through ByteBuffer for certain bytes and word with sectionsJava NIO 通过 ByteBuffer 扫描某些字节和带有节的字
【发布时间】:2018-01-05 18:02:47
【问题描述】:

好的,所以我正在尝试做一些看起来应该相当简单的事情,但是有了这些新的 NIO 接口,事情让我很困惑!这就是我想要做的,我需要扫描文件作为字节,直到遇到某些字节!当我遇到那些特定的字节时,需要抓取那段数据并对其进行处理,然后继续执行此操作。我原以为有了 ByteBuffer 中的所有这些标记、位置和限制,我就能做到这一点,但我似乎无法让它发挥作用!这是我到目前为止所拥有的......

测试文本:

this is a line of text a
this is line 2b
line 3
line 4
line etc.etc.etc.

Test.java:

import java.io.IOException;
import java.nio.ByteBuffer;
import java.nio.channels.FileChannel;
import java.nio.charset.Charset;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.nio.file.StandardOpenOption;

public class Test {
    public static final Charset ENCODING = Charset.forName("UTF-8");
    public static final byte[] NEWLINE_BYTE = {0x0A, 0x0D};

    public Test() {

        String pathString = "test.txt";

        //the path to the file
        Path path = Paths.get(pathString);

        try (FileChannel fc = FileChannel.open(path, 
                StandardOpenOption.READ, StandardOpenOption.WRITE, StandardOpenOption.CREATE)) {            
            if (fc.size() > 0) {
                int n;
                ByteBuffer buffer = ByteBuffer.allocate((int) fc.size());
                do {                    
                    n = fc.read(buffer);
                } while (n != -1 && buffer.hasRemaining());
                buffer.flip();
                int pos = 0;
                System.out.println("FILE LOADED: |" + new String(buffer.array(), ENCODING) + "|");
                do {
                    byte b = buffer.get();
                    if (b == NEWLINE_BYTE[0] || b == NEWLINE_BYTE[1]) {
                        System.out.println("POS: " + pos);
                        System.out.println("POSITION: " + buffer.position());
                        System.out.println("LENGTH: " + Integer.toString(buffer.position() - pos));
                        ByteBuffer lineBuffer = ByteBuffer.wrap(buffer.array(), pos + 1, buffer.position() - pos);
                        System.out.println("LINE: |" + new String(lineBuffer.array(), ENCODING) + "|");
                        pos = buffer.position();
                    }
                } while (buffer.hasRemaining());
            } 
        } catch (IOException ioe) {
           ioe.printStackTrace();
        }
    }
    public static void main(String args[]) {
        Test t = new Test();
    }
}

所以第一部分正在运行,fc.read(buffer) 函数只运行一次并将整个文件拉入 ByteBuffer。然后在第二个 do 循环中,我可以逐个字节地循环,当它遇到 \n(或 \r)时,它确实会命中 if 语句,但是我不知道如何得到它我刚刚查看了要使用的单独字节数组的部分字节!我已经尝试过拼接和各种翻转,我已经尝试过如上面的代码所示的 wrap,但似乎无法让它工作,两个缓冲区总是有完整的文件,所以我拼接或包裹它的任何东西!

我只需要逐字节循环文件,一次查看某个部分,然后我的最终目标,当我查看并找到正确的位置时,我想插入一些数据到正确的位置!我需要在“LINE:”处输出的 lineBuffer 只包含到目前为止我循环的部分字节!帮助,谢谢!

【问题讨论】:

  • TL;DR 有一个 ByteBuffer#wrap(byte[], int, int) 好像这就是你要找的东西
  • @Eugene 是这样的:ByteBuffer lineBuffer = ByteBuffer.wrap(buffer.array(), startOfLine, buffer.position());
  • 是的,看起来像......让我知道这是否有效
  • 仍然无法使其工作,我尝试创建一个包装在该部分之外的缓冲区。代码正在运行,但每次整个文件都在缓冲区中,而不仅仅是第一行!编辑问题以添加更新的代码。
  • 奇怪...我真的不想调试你的代码,但看看这个,因为它工作得很好:String test = "123456789"; ByteBuffer newB = ByteBuffer.wrap(test.getBytes(), 1, 3); System.out.println(StandardCharsets.UTF_8.decode(newB)); // 234

标签: java nio bytebuffer filechannel


【解决方案1】:

撇开 I/O 不谈,一旦您在 ByteBuffer 中有内容,通过 asCharBuffer() 将其转换为 CharBuffer 会简单得多。然后CharBuffer 实现CharSequence,它为您提供了很多String 和正则表达式方法可供使用。

【讨论】:

  • 我实际上打算将每一行显式转换为一个字符串并使用正则表达式对其进行解析,但我想在让 java 的解析器使用之前先在二进制文件中进行初始定义以验证一些事情将所有字符串转换为 UTF-16。
【解决方案2】:

这是我最终得到的解决方案,每次使用 ByteBuffer 的批量相对 get 函数来获取块。我想我正在使用 mark() 功能,尽管我使用了一个额外的变量 (pos) 来跟踪标记,因为我在 ByteBuffer 中找不到一个函数来返回标记本身的相对位置。另外,我有明确的功能可以按顺序查找 \r、\n 或两者。请记住,此代码仅适用于 UTF-8 编码数据。我希望这对其他人有帮助。

public class Test {
    public static final Charset ENCODING = Charset.forName("UTF-8");
    public static final byte[] NEWLINE_BYTES = {0x0A, 0x0D};

    public Test() {
        //test text file sequence of any strings followed by newline
        String pathString = "test.txt";
        Path path = Paths.get(pathString);

        try (FileChannel fc = FileChannel.open(path, 
                StandardOpenOption.READ, StandardOpenOption.WRITE, StandardOpenOption.CREATE)) {

            if (fc.size() > 0) {
                int n;
                ByteBuffer buffer = ByteBuffer.allocate((int) fc.size());
                do {                    
                    n = fc.read(buffer);
                } while (n != -1 && buffer.hasRemaining());
                buffer.flip();
                int newlineByteCount = 0;
                buffer.mark();
                do {
                    //get one byte at a time
                    byte b = buffer.get();

                    if (b == NEWLINE_BYTES[0] || b == NEWLINE_BYTES[1]) {
                        newlineByteCount++;

                        byte nextByte = buffer.get();
                        if (nextByte == NEWLINE_BYTES[1]) {
                            newlineByteCount++;
                        } else {
                            buffer.position(buffer.position() - 1);
                        }

                        int pos = buffer.position();
                        //reset the buffer back to the mark() position
                        buffer.reset();
                        //create an array just the right length and get the bytes we just measured out 
                        int length = pos - buffer.position() - newlineByteCount;
                        byte[] lineBytes = new byte[length];
                        buffer.get(lineBytes, 0, length);

                        String lineString = new String(lineBytes, ENCODING);
                        System.out.println("LINE: " + lineString);

                        buffer.position(buffer.position() + newlineByteCount);

                        buffer.mark();
                        newlineByteCount = 0;
                    } else if (newlineByteCount > 0) {

                    }
                } while (buffer.hasRemaining());
            } 
        } catch (IOException ioe) { ioe.printStackTrace(); }
    }
    public static void main(String args[]) { new Test(); }
}

【讨论】:

  • 查看我的答案以获得更通用的解决方案。
【解决方案3】:

我需要一些类似但比拆分单个缓冲区更通用的东西。就我而言,我有多个缓冲区;其实我的代码是对SpringStringDecoder的修改,可以将Flux<DataBuffer>(DataBuffer)转换成Flux<String>。

https://stackoverflow.com/a/48111196/839733

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-05-15
    • 1970-01-01
    • 1970-01-01
    • 2013-08-15
    • 2014-03-16
    • 1970-01-01
    相关资源
    最近更新 更多