【问题标题】:How to replace unicode characters by ascii characters in Python (perl script given)?如何在 Python 中用 ascii 字符替换 unicode 字符(给定的 perl 脚本)?
【发布时间】:2011-02-11 15:39:49
【问题描述】:

我正在尝试学习 python,但不知道如何将以下 perl 脚本翻译成 python:

#!/usr/bin/perl -w                     

use open qw(:std :utf8);

while(<>) {
  s/\x{00E4}/ae/;
  s/\x{00F6}/oe/;
  s/\x{00FC}/ue/;
  print;
}

脚本只是将 unicode 变音符号更改为替代 ascii 输出。 (所以完整的输出是 ascii。)我会很感激任何提示。谢谢!

【问题讨论】:

标签: python perl unicode diacritics


【解决方案1】:
  • 使用fileinput 模块循环标准输入或文件列表,
  • 将您从 UTF-8 读取的行解码为 un​​icode 对象
  • 然后使用 translate 方法映射您想要的任何 unicode 字符

translit.py 看起来像这样:

#!/usr/bin/env python2.6
# -*- coding: utf-8 -*-

import fileinput

table = {
          0xe4: u'ae',
          ord(u'ö'): u'oe',
          ord(u'ü'): u'ue',
          ord(u'ß'): None,
        }

for line in fileinput.input():
    s = line.decode('utf8')
    print s.translate(table), 

你可以这样使用它:

$ cat utf8.txt 
sömé täßt
sömé täßt
sömé täßt

$ ./translit.py utf8.txt 
soemé taet
soemé taet
soemé taet
  • 更新:

如果您使用的是 python 3 字符串,默认情况下是 unicode,如果它包含非 ASCII 字符甚至非拉丁字符,则不需要对其进行编码。所以解决方案如下:

line = 'Verhältnismäßigkeit, Möglichkeit'

table = {
         ord('ä'): 'ae',
         ord('ö'): 'oe',
         ord('ü'): 'ue',
         ord('ß'): 'ss',
       }

line.translate(table)

>>> 'Verhaeltnismaessigkeit, Moeglichkeit'

【讨论】:

  • 我猜最后一行应该是 print s.translate(table).encode('ascii', 'ignore') 来获得 ascii 输出。
  • 目标似乎是消除德语文本的变音,使其易于理解。这段代码中ord(u'ß'): None 的作用是删除 ß(“eszett”)字符。应该是ord(u'ß'): u'ss'。点赞??接受的答案???
  • 哦。来。在。我试图展示地图的不同可能性。
  • 你选择了一个非常糟糕的例子来说明如何做一些 OP 没有表明他想要或需要的事情。
  • @john:如果您将 OP 的问题与他上面的评论('ignore')逐字逐句地结合起来,它将具有 exact same结果,所以不要吹毛求疵了。
【解决方案2】:

要转换为 ASCII,您可能需要尝试 ASCII, Dammit 或 this recipe,归结为:

>>> title = u"Klüft skräms inför på fédéral électoral große"
>>> import unicodedata
>>> unicodedata.normalize('NFKD', title).encode('ascii','ignore')
'Kluft skrams infor pa federal electoral groe'

【讨论】:

  • 这根本不是原始 .pl 所做的(主要是正确音译德语特殊字符)
  • 从德语变音符号中去除点与从“x”中去除一条腿并写“y”或将“d”替换为“b”一样有意义,因为“有点像” .
  • 不,您可能会遇到冲突,因为您将不同的字符串映射到同一个字符串。
【解决方案3】:

我用translitcodec

>>> import translitcodec
>>> print '\xe4'.decode('latin-1')
ä
>>> print '\xe4'.decode('latin-1').encode('translit/long').encode('ascii')
ae
>>> print '\xe4'.decode('latin-1').encode('translit/short').encode('ascii')
a

您可以将解码语言更改为您需要的任何语言。您可能需要一个简单的函数来减少单个实现的长度。

def fancy2ascii(s):
    return s.decode('latin-1').encode('translit/long').encode('ascii')

【讨论】:

    【解决方案4】:

    您可以尝试unidecode 将 Unicode 转换为 ascii,而不是手动编写正则表达式。它是Text::Unidecode Perl 模块的Python 端口:

    #!/usr/bin/env python
    import fileinput
    import locale
    from contextlib import closing
    from unidecode import unidecode # $ pip install unidecode
    
    def toascii(files=None, encoding=None, bufsize=-1):
        if encoding is None:
            encoding = locale.getpreferredencoding(False)
        with closing(fileinput.FileInput(files=files, bufsize=bufsize)) as file:
            for line in file: 
                print unidecode(line.decode(encoding)),
    
    if __name__ == "__main__":
        import sys
        toascii(encoding=sys.argv.pop(1) if len(sys.argv) > 1 else None)
    

    它使用FileInput 类来避免全局状态。

    例子:

    $ echo 'äöüß' | python toascii.py utf-8
    aouss
    

    【讨论】:

      【解决方案5】:

      又快又脏(python2):

      def make_ascii(string):
          return string.decode('utf-8').replace(u'ü','ue').replace(u'ö','oe').replace(u'ä','ae').replace(u'ß','ss').encode('ascii','ignore');
      

      【讨论】:

        猜你喜欢
        • 2018-07-02
        • 1970-01-01
        • 1970-01-01
        • 2016-06-20
        • 2016-07-04
        • 2015-02-27
        • 1970-01-01
        • 2014-12-03
        • 1970-01-01
        相关资源
        最近更新 更多