【问题标题】:Finding the equivalent Unicode codepoint given an extended ASCII codepoint and a codepage in Java?给定扩展的 ASCII 代码点和 Java 中的代码页,找到等效的 Unicode 代码点?
【发布时间】:2020-05-20 23:23:56
【问题描述】:

我正在尝试编写一种方法来查找给定特定代码页的 ASCII 中相同视觉字符的 Unicode 中的等效代码点

例如,给定一个字符 char c = 128,它在 Windows-1252 代码页中是“€”,运行该方法

int result = asUnicode(c, "windows-1252")

应该给出8364 或相同的char c = 128,在Windows-1251 代码页中是“Ђ”,运行该方法

int result = asUnicode(c, "windows-1251")

应该给1026

如何在 Java 中做到这一点?

【问题讨论】:

    标签: java unicode character-encoding ascii


    【解决方案1】:

    c 不应该是真正的char,而是相应编码中字节的byte[],例如。 windows-1252.

    对于这种简单的情况,我们可以自己将char 包装成byte[]。

    您需要将这些字节解码为代表 BMP 代码点的 Java 的 char 类型。然后你返回对应的。

    public static int asUnicode(char c, String charset) throws Exception {
        CharBuffer result = Charset.forName(charset).decode(ByteBuffer.wrap(new byte[] { (byte) c }));
        int unicode;
        char first = result.get();
        if (Character.isSurrogate(first)) {
            unicode = Character.toCodePoint(first, result.get());
        } else {
            unicode = first;
        }
        return unicode;
    }
    

    以下

    public static void main(String[] args) throws Exception {
        char c = 128;
        System.out.println(asUnicode(c, "windows-1252"));
        System.out.println(asUnicode(c, "windows-1251"));
    }
    

    打印

    8364
    1026
    

    【讨论】:

    • 谢谢,答案很快!不知道为什么它必须在 byte[] 但它有效。
    • @user1589188 看起来 windows-1251 和 windows-1252 都是单字节编码。如果它需要超过 2 个字节来表示其字符,那么您需要的不仅仅是 char。此外,Charset API 消耗字节(在解码期间)。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-07-04
    • 1970-01-01
    相关资源
    最近更新 更多