【问题标题】:How do I get this to encode properly?我怎样才能让它正确编码?
【发布时间】:2013-07-05 20:14:16
【问题描述】:

我有一个带有俄语文本的 XML 文件:

<p>все чашки имеют стандартный посадочный диаметр - 22,2 мм</p>

我使用xml.etree.ElementTree 以各种方式对其进行操作(从未接触过文本内容)。然后,我使用ElementTree.tostring:

info["table"] = ET.tostring(table, encoding="utf8") #table is an Element

然后我用这个字符串做一些其他的事情,最后把它写到一个文件中

f = open(newname, "w")
output = page_template.format(**info)
f.write(output)
f.close()

我在我的文件中结束了这个:

<p>\xd0\xb2\xd1\x81\xd0\xb5 \xd1\x87\xd0\xb0\xd1\x88\xd0\xba\xd0\xb8 \xd0\xb8\xd0\xbc\xd0\xb5\xd1\x8e\xd1\x82 \xd1\x81\xd1\x82\xd0\xb0\xd0\xbd\xd0\xb4\xd0\xb0\xd1\x80\xd1\x82\xd0\xbd\xd1\x8b\xd0\xb9 \xd0\xbf\xd0\xbe\xd1\x81\xd0\xb0\xd0\xb4\xd0\xbe\xd1\x87\xd0\xbd\xd1\x8b\xd0\xb9 \xd0\xb4\xd0\xb8\xd0\xb0\xd0\xbc\xd0\xb5\xd1\x82\xd1\x80 - 22,2 \xd0\xbc\xd0\xbc</p>

如何正确编码?

【问题讨论】:

  • 您的意思是文件中有文字反斜杠?或者这是 python bytes 对象的表示?那是正确的。
  • @mata:这就是我文件中的字面意思。

标签: python xml encoding utf-8 python-3.x


【解决方案1】:

你使用

 info["table"] = ET.tostring(table, encoding="utf8")

返回bytes。然后稍后将其应用于格式字符串,即str(unicode),如果这样做,您最终将得到字节对象的表示。

如果你使用,etree 可以返回一个 unicode 对象:

 info["table"] = ET.tostring(table, encoding="unicode")

【讨论】:

    【解决方案2】:

    问题在于 ElementTree.tostring 返回的是二进制对象,而不是实际的字符串。答案是:

    info["table"] = ET.tostring(table, encoding="utf8").decode("utf8")
    

    【讨论】:

      【解决方案3】:

      试试这个 - 输出参数只是没有 utf-8 编码的俄语字符串。

      import codecs
      #output=u'все чашки имеют стандартный посадочный диаметр'
      with codecs.open(newname, "w", "utf-16") as stream: #or utf-8
          stream.write(output + u"\n")
      

      【讨论】:

      • 它是 python3,字符串不需要 u 前缀。它甚至会导致 3.3 之前的版本出错。
      猜你喜欢
      • 2021-05-27
      • 1970-01-01
      • 1970-01-01
      • 2012-08-02
      • 2021-07-06
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-01-19
      相关资源
      最近更新 更多