【发布时间】:2013-10-17 12:41:20
【问题描述】:
基本上在我的学生数据中,我遇到了一个问题,如您所见,我的数据中出现了奇怪的 sumbols:MAIN £1.00when 它应该显示MAIN £1.00
下面是我的代码的 sn-p,它从网站上抓取某些学生信息以获得学生折扣,并最终将其写入文件。
# -*- coding: utf-8 -*-
totals = main.find_all('p')
for total in totals:
if total .find(text=re.compile("Main:")):
total = total.get_text()
if u"Main £" in total:
pull1 = re.search(r'(MAIN) (\D\w+\D\d+)', total)
pull2 = re.search(r'(MAINER) (\D\w+\D\d+)', total)
if pull1:
rpr_data.append(pull1.group(0).title())
print pull1.group(0).title()
if pull2:
rpr_data.append(pull2.group(0).title())
print pull2.group(0).title()
with open('RPR.txt','w') as rpr_file:
rpr_file.write('\n'.join(rpr_data).encode("UTF-8"))
当我尝试在脚本 Matching three variables from textfile to csv and writing variables to the csv on matched rows 中重新使用此数据时,即使文本文件中的数据在写入 CSV 时没有奇怪的 Â 符号,符号又回来了...
我怎样才能正确地永久消除这个Â 符号?
【问题讨论】:
-
首先,您确定脚本实际上保存为 UTF-8 文本文件,而不是 Latin-1/cp1252/etc。 (只是在顶部放一个编码声明注释并不会改变你的文本编辑器使用的编码,除了 emacs,它只是对 Python 说谎……)
-
另外,您在哪里看到
MAIN £1.00?打开RPR.txt后在Notepad.exe中?还是……? -
@abarnert 我在某些印刷品上看到了这一点。使用行
total = total.get_text()以某种方式剥离Â符号,但在另一个线程中运行匹配脚本后,我得到Â重新出现在列中 -
啊,那是因为你的终端的字符集是 Latin-1/cp1252/etc.,而 Python 正在打印 UTF-8。 (能否告诉我们您使用的是什么平台,如果不是 Windows,您使用的是什么终端,所以我不必一直猜测?)
标签: python csv python-2.7 utf-8 encode