【问题标题】:Python: calling upper() on words containing non-latin charactersPython:对包含非拉丁字符的单词调用 upper()
【发布时间】:2015-06-09 08:28:29
【问题描述】:

我有一个包含单词的文件,例如

А
б
Вв
Гг

(非拉丁字母) 等等

我想得到这个:

А
Б
ВВ
ГГ

代码运行后我看不到任何变化

代码如下:

f = open('sample.csv')
for line in f:
    for sampleword in line.split():
        print sampleword.upper()

非拉丁字符不大写。有什么问题?

【问题讨论】:

  • 输入、输出和预期输出是什么?
  • 它对我有用。它打印 A AB ABC。这不是你想要的吗?
  • 是的,没错。也许问题是我使用非英文字母?
  • words = ['ab', 'cd', 'ef'] 和 for w in words: print w.upper() 工作正常,无法重现。您能否验证您从文件中读取的具体内容是什么?
  • 那么,真正的问题是:如何将非英语/拉丁字母大写?你能相应地编辑你的问题吗?请添加示例!

标签: python python-2.x uppercase


【解决方案1】:

Python 2 中非拉丁字母大写的解决方案是使用 unicode 字符串:

words = [u'łuk', u'ćma']
assert [w.upper() for w in words] == [u'ŁUK', u'ĆMA']

要从文件中读取 unicode,您可以参考official Python manual:

因此从文件中读取 Unicode 很简单:

import codecs
f = codecs.open('unicode.rst', encoding='utf-8')
for line in f:
    print repr(line)

【讨论】:

  • 但是文件呢?
  • 好的,这很好,但现在我的输出文件中有 u'\u0410\u041a\u0418\u0411\u0410\u041d\u041a' 之类的东西
  • 你在做print repr(word.upper())吗? repr() 仅用于表明字符串是unicode,而不是str。您可以简单地打印行本身 - print word.upper().
  • 否则我会收到错误消息 UnicodeDecodeError: 'ascii' codec can't decode byte 0xd0 in position 0: ordinal not in range(128)
  • @PeterKungurtsev:您的输入文件编码是什么?也许最初不是 UTF-8?
猜你喜欢
  • 2016-05-19
  • 2017-04-28
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2012-09-20
  • 1970-01-01
  • 2013-01-20
  • 2015-06-26
相关资源
最近更新 更多