【问题标题】:how to parse json from stream in java如何从java中的流中解析json
【发布时间】:2015-05-16 16:13:52
【问题描述】:

我需要编写一个使用套接字与服务器通信的客户端。消息协议为 json 格式。服务器将根据需要将多个 json 块推送到客户端。 消息是这样的:

{"a": 1, "b": { "c": 1}}{"a": 1, "b": { "c": 1}}... 

您可以看到json块之间没有分隔符或标识符。

我能找到的 json 解析器(如 fastjson、jackson)都只能将流作为一个完整的 json 块来处理,即使是它们提供的流 api。当我使用这些api解析流时,它们会在第一个json块的末尾抛出一个异常,表示下一个标记“{”无效。

java中是否有一个json解析器可以处理我的问题?还是有其他方法可以解决这个问题?

【问题讨论】:

  • 我检查了上面链接中的答案。情况有点不同。服务器通过流提供json,所以我不能轻易使用正则表达式。
  • 你不能将流转换为字符串然后使用正则表达式吗?
  • 所以结果就像一个数组,只是它不是任何 JSON 解析器的数组。你有嵌套的 JSON 块吗?我的意思是,你能假设}{ 总是将一个元素与下一个元素分开吗?
  • @Magnamag 是的,我有嵌套块。所以我不能只用“}{”分隔块,字段值中也可能有“{”、“}”。

标签: java json stream


【解决方案1】:

最后它在我的情况下没有在 java 中缝合任何 JSON 解析器。 我正在使用 netty 构建网络应用程序。对于 nio,当有来自网络的数据时,调用 ByteToMessageDecoder 中的 decode 方法。 在这种方法中,我需要从 ByteBuf 中找出 JSON 块。

由于没有可用的 JSON 解析器,我编写了一个方法来从 ByteBuf 中拆分 JSON 块。

    public static void extractJsonBlocks(ByteBuf buf, List<Object> out) throws UnsupportedEncodingException {
    // the total bytes that can read from ByteBuf
    int readable = buf.readableBytes();
    int bracketDepth = 0;
    // when found a json block, this value will be set
    int offset = 0;
    // whether current character is in a string value
    boolean inStr = false;
    // a temporary bytes buf for store json block
    byte[] data = new byte[readable];
    // loop all the coming data
    for (int i = 0; i < readable; i++) {
        // read from ByteBuf
        byte b = buf.readByte();
        // put it in the buffer, be care of the offset
        data[i - offset] = b;
        if (b == SYM_L_BRACKET && !inStr) {
            // if it a left bracket and not in a string value
            bracketDepth++;
        } else if (b == SYM_R_BRACKET && !inStr) {
            // if it a right bracket and not in a string value
            if (bracketDepth == 1) {
                // if current bracket depth is 1, means found a whole json block
                out.add(new String(data, "utf-8").trim());
                // create a new buffer
                data = new byte[readable - offset];
                // update the offset
                offset = i;
                // reset the bracket depth
                bracketDepth = 0;
            } else {
                bracketDepth--;
            }
        } else if (b == SYM_QUOTE) {
            // when find a quote, we need see whether preview character is escape.
            byte prev = i == 0 ? 0 : data[i - 1 - offset];
            if (prev != SYM_ESCAPE) {
                inStr = !inStr;
            }
        }

    }
    // finally there may still be some data left in the ByteBuf, that can not form a json block, they should be used to combine with the following datas
    // so we need to reset the reader index to the first byte of the left data
    // and discard the data used for json blocks
    buf.readerIndex(offset == 0 ? offset : offset + 1);
    buf.discardReadBytes();
}

也许这不是一个完美的解析器,但它现在适用于我的应用程序。

【讨论】:

    【解决方案2】:

    您可以使用Genson 完成此操作。 首先将其配置为允许“允许”解析,然后反序列化迭代器中的值。当您在迭代器上调用 next 时,Genson 将一一解析对象。因此,您可以通过这种方式解析非常大的输入。

    Genson genson = new GensonBuilder().usePermissiveParsing(true).create();
    ObjectReader reader = genson.createReader(inputStream);
    Iterator<SomeObject> iterator = genson.deserializeValues(reader, GenericType.of(SomeObject.class));
    

    这部分 API 有点冗长,因为用例并不常见。

    更新 在 Genson 1.4 中,usePermissiveParsing 已被删除,以支持默认情况下接受未包装在数组中的根值。见https://github.com/owlike/genson/issues/78

    【讨论】:

    • 如果我能在 GensonBuilder 上找到这个 usePermissiveParsing 方法就该死
    • 从 genson 1.4 开始 usePermissiveParsing 已被删除,默认情况下它应该可以工作。见github.com/owlike/genson/issues/78
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-02-20
    • 2022-08-18
    • 2017-04-27
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多