【问题标题】:How can i replace every emoji in a string with their unicode in java?java - 如何用java中的unicode替换字符串中的每个表情符号?
【发布时间】:2020-04-10 02:30:38
【问题描述】:

我有一个这样的字符串:

"\"title\":\"????TEST title value ????\",\"text\":\"???? TEST text value.\"" ...

我想用它们的 unicode 值替换每个表情符号,如下所示:

"\"title\":\"U+1F47ATEST title value U+1F601\",\"text\":\"U+1F496 TEST text value.\"" ...

在网上搜索了很多之后,我找到了一种使用以下代码将一个符号“翻译”为它们的 unicode 的方法:

String s = "????";
int emoji = Character.codePointAt(s, 0); 
String unumber = "U+" + Integer.toHexString(emoji).toUpperCase();

但是现在如何更改我的代码以获取字符串中的所有表情符号?

附:它可以是 \Uxxxxx 或 U+xxxxx 格式

【问题讨论】:

  • 如果您的实际文本包含简单的纯文本字符串 U+12345 怎么办?或者你能保证这永远不会发生吗?

标签: java unicode emoji


【解决方案1】:

您无需在代码中指定任何代码点范围,也无需担心代理项。相反,只需指定您希望将其字符显示为 Unicode 转义符的 Unicode 块。这是通过使用Character.UnicodeBlock 类中的字段声明来实现的。例如,判断?(0x1F601)是否为表情符号:

boolean emoticon = Character.UnicodeBlock.EMOTICONS.equals(Character.UnicodeBlock.of("?".codePointAt(0)));
System.out.println("Is ? an emoticon? " + emoticon); // Prints true.

这是通用代码。它将处理任何String,如果在指定的 Unicode 代码块中定义了单个字符,则将它们显示为它们的 Unicode 等价物:

package symbolstounicode;

import java.util.List;
import java.util.stream.Collectors;

public class SymbolsToUnicode {

    public static void main(String[] args) {

        Character.UnicodeBlock[] blocksToConvert = new Character.UnicodeBlock[]{
            Character.UnicodeBlock.EMOTICONS, 
            Character.UnicodeBlock.MISCELLANEOUS_SYMBOLS_AND_PICTOGRAPHS};
        String input = "\"title\":\"?TEST title value ?\",\"text\":\"? TEST text value.\"";
        String output = SymbolsToUnicode.toUnicode(input, blocksToConvert);

        System.out.println("String to convert: " + input);
        System.out.println("Converted string: " + output);
        assert ("\"title\":\"U+1F47ATEST title value U+1F601\",\"text\":\"U+1F496 TEST text value.\"".equals(output));
    }

    // Converts characters in the supplied string found in the specified list of UnicodeBlocks to their Unicode equivalents.
    static String toUnicode(String s, final Character.UnicodeBlock[] blocks) {

        StringBuilder sb = new StringBuilder("");
        List<Integer> cpList = s.codePoints().boxed().collect(Collectors.toList());

        cpList.forEach(cp -> sb.append(SymbolsToUnicode.inCodeBlock(cp, blocks) ? 
                "U+" + Integer.toHexString(cp).toUpperCase() : Character.toString(cp)));
        return sb.toString();
    }

    // Returns true if the supplied code point is within one of the specified UnicodeBlocks.
    static boolean inCodeBlock(final int cp, final Character.UnicodeBlock[] blocksToConvert) {

        for (Character.UnicodeBlock b : blocksToConvert) {
            if (b.equals(Character.UnicodeBlock.of(cp))) {
                return true;
            }
        }
        return false;
    }
}

这是输出,使用 OP 中的测试数据:

run:
String to convert: "title":"?TEST title value ?","text":"? TEST text value."
Converted string: "title":"U+1F47ATEST title value U+1F601","text":"U+1F496 TEST text value."
BUILD SUCCESSFUL (total time: 0 seconds)

注意事项:

  • 我使用字体Segoe UI Symbol作为代码和输出窗口来正确渲染符号。
  • 代码中的基本思想是:
    • 首先,指定要转换的String,以及字符应转换为Unicode 的Unicode 代码块。
    • 接下来,使用String.codePoints() 将String 转换为一组代码点,并将它们存储在List 中。
    • 最后,对于每个代码点,确定它是否存在于任何指定的 Unicode 块中,并在必要时进行转换。

【讨论】:

    【解决方案2】:

    表情符号分散在不同的unicode blocks 中。例如?(0x1F47A) 和?(0x1F496) 来自Miscellaneous Symbols and Pictographs,而?(0x1F601) 来自Emoticons

    如果您想过滤掉符号,您需要决定要使用哪些 unicode 块(或其范围)。例如:

        String s = "\"title\":\"?TEST title value ?\",\"text\":\"? TEST text value.\"";
        StringBuilder sb = new StringBuilder();
        for (int i = 0, l = s.length() ; i < l ; i++) {
          char ch = s.charAt(i);
          if (Character.isHighSurrogate(ch)) {
            i++;
            char ch2 = s.charAt(i); // Load low surrogate
            int codePoint = Character.toCodePoint(ch, ch2);
            if ((codePoint >= 0x1F300) && (codePoint <= 0x1F64F)) { // Miscellaneous Symbols and Pictographs + Emoticons
              sb.append("U+").append(Integer.toHexString(codePoint).toUpperCase());
            } else { // otherwise just add characters as is
              sb.append(ch);
              sb.append(ch2);
            }
          } else { // if not a surrogate, just add the character
            sb.append(ch);
          }
        }
        String result = sb.toString();
        System.out.println(result); // "title":"U+1F47ATEST title value U+1F601","text":"U+1F496 TEST text value."
    

    要仅获取表情符号,您可以使用 this list 等来缩小条件范围

    但如果你想转义任何代理符号,你可以去掉codePoint检查代码内部

    【讨论】:

    • 据我所知,在任何给定的 unicode 平面中都没有包含所有表情符号且不与实际字母相交的单一范围。因此,您的答案范围不正确,唯一正确的检查方法是下载并使用 unicode-emoji 文件。
    • @M.Prokhorov 是的。这只是一个例子。
    • 效果很好,但我不确定我会收到哪些 unicode 块,我可以将 0x1F300 和 0x1F64F 与第一个和最后一个 unicode 块交换吗?如果是,我在哪里可以找到它们?
    • @BladeITA,你可以看这里:unicode.org/charts - 这些都是今天的标准块,你需要找到你想要替换的块。请注意,“正常”范围也有类似表情符号的字符,在“Dingbats”表中。
    【解决方案3】:

    试试这个解决方案:

    String s = "your string with emoji";
    
    StringBuilder sb = new StringBuilder();
    
    for (int i = 0; i < s.length(); i++) {
      if (Character.isSurrogate(s.charAt(i))) {
        Integer res = Character.codePointAt(s, i);
        i++;
        sb.append("U+" + Integer.toHexString(res).toUpperCase());
      } else {
        sb.append(s.charAt(i));
      }
    }
    
    //result
    System.out.println(sb.toString());
    

    【讨论】:

    • 是的,但其他代理字符可能会导致同样的问题。如果我们只有表情符号作为特定字符,解决方案将正确处理。
    • 这里没有“问题”。他想用他们的 unicode refs 替换表情符号。据我所知 - 只是表情符号,而不是每个“不寻常”的字符。
    • 我试图复制这段代码,但我得到了错误:'方法 isSurrogate(char) 未定义字符类型',我该如何解决?
    • 尝试导入java.lang.Character;你有这样的课吗?
    • 没有,但即使添加了它,我仍然遇到同样的问题。 AterLux 的解决方案很有效
    猜你喜欢
    • 2014-01-12
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-04-22
    • 2016-03-22
    • 2020-05-12
    • 2011-11-11
    相关资源
    最近更新 更多