【问题标题】:Convert string representation of a hexadecimal byte array to a string with non ascii characters in Java在 Java 中将十六进制字节数组的字符串表示形式转换为具有非 ascii 字符的字符串
【发布时间】:2020-07-27 15:18:31
【问题描述】:

我有一个字符串被客户端发送到请求负载中:

"[0xc3][0xa1][0xc3][0xa9][0xc3][0xad][0xc3][0xb3][0xc3][0xba][0xc3][0x81][0xc3][0x89][0xc3][0x8d][0xc3][0x93][0xc3][0x9a]Departms"

我想要一个字符串 "áéíóúÁÉÍÓÚDepartms"。我如何在 Java 中做到这一点?

问题是我无法控制客户端编码此字符串的方式。似乎客户端只是以这种格式编码非 ascii 字符并按原样发送 ascii 字符(请参阅最后的“Departms”)。

【问题讨论】:

    标签: java hex non-ascii-characters string-decoding


    【解决方案1】:

    方括号内的内容似乎是用 UTF-8 编码的字符,但以一种奇怪的方式转换为十六进制字符串。您可以做的是找到每个看起来像[0xc3] 的实例并将其转换为相应的字节,然后从字节中创建一个新的字符串。

    不幸的是,没有很好的工具来处理字节数组。这是一个快速而肮脏的解决方案,它使用正则表达式查找这些十六进制代码并将其替换为 latin-1 中的相应字符,然后通过重新解释字节来修复它。

    String bracketDecode(String str) {
        Pattern p = Pattern.compile("\\[(0x[0-9a-f]{2})\\]");
        Matcher m = p.matcher(str);
        StringBuilder sb = new StringBuilder();
        while (m.find()) {
            String group = m.group(1);
            Integer decode = Integer.decode(group);
            // assume latin-1 encoding
            m.appendReplacement(sb, Character.toString(decode));
        }
        m.appendTail(sb);
        // oh no, latin1 is not correct! re-interpret bytes in utf-8
        byte[] bytes = sb.toString().getBytes(StandardCharsets.ISO_8859_1);
        return new String(bytes, StandardCharsets.UTF_8);
    }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-03-11
      • 2015-05-26
      • 2011-10-26
      • 2013-01-14
      • 1970-01-01
      • 2021-10-31
      相关资源
      最近更新 更多