【发布时间】:2015-10-08 21:24:36
【问题描述】:
首先,我看过这里:Sublime Text 3, Python 3 and UTF-8 don't like each other 并阅读了The Absolute Minimum Every Software Developer Absolutely, Positively Must Know About Unicode and Character Sets,但我仍然不知道以下内容:
从在 Sublime 中创建的文件运行 Python(不是编译)并在 XP 机器上通过命令提示符执行
我有几个以重音命名的文本文件(主要是德语、西班牙语和法语)。我想删除重音字符(变音符号、锐音符、坟墓、cidillas 等)并将它们替换为它们的等价非重音外观。
如果重音是脚本中的字符串,我可以去掉它们。但是访问同名的文本文件会导致 strippAcent 函数失败。我完全没有想法,因为我认为这是由于与 Sublime 和 Python 的冲突。
这是我的脚本
# -*- coding: utf-8 -*-
import unicodedata
import os
def stripAccents(s):
try:
us = unicode(s,"utf-8")
nice = unicodedata.normalize("NFD", us).encode("ascii", "ignore")
print nice
return nice
except:
print ("Fail! : %s" %(s))
return None
stripAccents("Découvrez tous les logiciels à télécharger")
# Decouvrez tous les logiciels a telecharger
stripAccents("Östblocket")
# Ostblocket
stripAccents("Blühende Landschaften")
# Bluhende Landschaften
root = "D:\\temp\\test\\"
for path, subdirs, files in os.walk(root):
for name in files:
x = name
x = stripAccents(x)
记录在案:
C:\chcp
给我 437
完整的错误是:
C:\WINDOWS\system32>D:\LearnPython\unicode_accents.py
Decouvrez tous les logiciels a telecharger
Ostblocket
Bluhende Landschaften
Traceback (most recent call last):
File "D:\LearnPython\unicode_accents.py", line 37, in <module>
x = stripAccents(x)
File "D:\LearnPython\unicode_accents.py", line 8, in stripAccents
us = unicode(s,"utf-8")
UnicodeDecodeError: 'utf8' codec can't decode byte 0xfc in position 2: invalid start byte
C:\WINDOWS\system32>
【问题讨论】:
-
您到底想做什么 - 重命名文件?因为你没有在你的脚本中这样做。您只需阅读它们的名称,然后通过您的函数运行名称。您永远不会尝试将该文件名写回磁盘...
-
请包含运行脚本的完整输出,包括错误消息。
-
@MattDMo 是的,我现在正在调试。它目前没有重命名。在整理字符串之前,我没有包含重命名代码。
-
不要捕获异常并打印
Fail,而是注释掉except块(以及try:)并查看实际错误是什么。然后你就会知道如何解决它。 -
@MattDMo 按要求添加了错误。
标签: python-2.7 unicode sublimetext2 python-unicode