【问题标题】:How do I preserve utf-8 JSON values and write them correctly to a utf-8 txt file in Python 3.8.2?如何保留 utf-8 JSON 值并将它们正确写入 Python 3.8.2 中的 utf-8 txt 文件?
【发布时间】:2020-07-05 18:53:33
【问题描述】:

我最近编写了一个 python 脚本,用于从 JSON 文件中提取一些数据,并使用它为以下语句生成一些 SQL 插入值:

INSERT INTO `card`(`artist`,`class_pid`,`collectible`,`cost`, `dbfid`, `api_db_id`, `name`, `rarity`, `cardset_pid`, `cardtype`, `attack`, `health`, `race`, `durability`, `armor`,`multiclassgroup`, `text`) VALUES ("generated entry goes here")

我的 SQL 表中某些属性的名称不同,但使用的值相同(JSON 文件/Python 脚本中的示例 cardClass 在 SQL 表中称为 class_pid)。从脚本生成的值是有效的 SQL,并且可以成功插入到数据库中,但是我注意到在生成的 export.txt 文件中,一些值与原来的值不同。例如,来自 utf-8 编码 JSON 文件的以下 JSON 条目:

[{"artist":"Arthur Bozonnet","attack":3,"cardClass":8,"collectible":1,"cost":2,"dbfId":2545,"flavor":"And he can't get up.","health":2,"id":"AT_003","mechanics":["HEROPOWER_DAMAGE"],"name":"Fallen Hero","rarity":"RARE","set":1,"text":"Your Hero Power deals 1 extra damage.","type":"MINION"},{"artist":"Ivan Fomin","attack":2,"cardClass":11,"collectible":1,"cost":2,"dbfId":54256,"flavor":"Were you expectorating another bad pun?","health":4,"id":"ULD_182","mechanics":["TRIGGER_VISUAL"],"name":"Spitting Camel","race":"BEAST","rarity":"COMMON","set":22,"text":"[x]At the end of your turn,\n  deal 1 damage to another  \nrandom friendly minion.","type":"MINION"}]

产生这个输出:

('Arthur Bozonnet',8,1,2,'2545','AT_003','Fallen Hero','RARE',1,'MINION',3,2,'NULL',0,0,'NULL','Your Hero Power deals 1\xa0extra damage.'),('Ivan Fomin',11,1,2,'54256','ULD_182','Spitting Camel','COMMON',22,'MINION',2,4,'BEAST',0,0,'NULL','[x]At the end of your turn,\n\xa0\xa0deal 1 damage to another\xa0\xa0\nrandom friendly minion.')

如您所见,JSON 条目中的某些值已以某种方式更改,就好像在某处更改了文本编码一样,即使在我的脚本中我确保 JSON 文件是使用 utf-8 编码打开的,并且生成的文本文件也被打开并写入 utf-8 以匹配 JSON 文件。我的目标是完全保留 JSON 文件中的值,并将这些值传输到生成的 SQL 值条目中,就像它们在 JSON 中一样。例如,在生成的 SQL 中,我希望第二个条目的“文本”值为:

"[x]At the end of your turn,\n  deal 1 damage to another  \nrandom friendly minion."

代替:

"[x]At the end of your turn,\n\xa0\xa0deal 1 damage to another\xa0\xa0\nrandom friendly minion."

我尝试使用诸如 unicodedata.normalize() 之类的函数,但不幸的是它似乎并没有以任何方式改变输出。 这是我为生成 SQL 值而编写的脚本:

import json
import io

chosen_keys = ['artist','cardClass','collectible','cost',
'dbfId','id','name','rarity','set','type','attack','health',
'race','durability','armor',
'multiClassGroup','text']

defaults = ['NULL','0','0','0',
'NULL','NULL','NULL','NULL','0','NULL','0','0',
'NULL','0','0',
'NULL','NULL']

def saveChangesString(dataList, filename):
  with io.open(filename, 'w', encoding='utf-8') as f:
    f.write(dataList)
    f.close()

def generateSQL(json_dict):
    count = 0
    endCount = 1
    records = ""
    finalState = ""
    print('\n'+str(len(json_dict))+' records will be processed\n')
    for i in json_dict:
        entry = "("
        jcount = 0
        for j in chosen_keys:
            if j in i.keys():
                if str(i.get(j)).isdigit() and j != 'dbfId':
                    entry = entry + str(i.get(j))
                else:
                    entry = entry + repr(str(i.get(j)))
            else:
                if str(defaults[jcount]).isdigit() and j != 'dbfId':
                    entry = entry + str(defaults[jcount])
                else:
                    entry = entry + repr(str(defaults[jcount]))
            if jcount != len(chosen_keys)-1:
                entry = entry+","
            jcount = jcount + 1
        entry = entry + ")"
        if count != len(json_dict)-1:
                entry = entry+","
        count = count + 1
        if endCount % 100 == 0 and endCount >= 100 and endCount < len(json_dict):
            print('processed records '+str(endCount - 99)+' - '+str(endCount))
            if endCount + 100 > len(json_dict):
                finalState = 'processed records '+str(endCount+1)+' - '+str(len(json_dict))
        if endCount == len(json_dict):
            print(finalState)
        records = records + entry
        endCount = endCount + 1
    saveChangesString(records,'export.txt')
    print('done')
                
with io.open('cards.collectible.sample.example.json', 'r', encoding='utf-8') as f:
    json_to_dict = json.load(f)
    f.close()

generateSQL(json_to_dict)

任何帮助都将不胜感激,因为我实际使用的 JSON 文件包含超过 2000 个条目,因此我宁愿避免手动编辑内容。谢谢。

另外SQL表结构代码为:

-- phpMyAdmin SQL Dump
CREATE TABLE `card` (
  `pid` int(10) NOT NULL,
  `api_db_id` varchar(50) NOT NULL,
  `dbfid` varchar(50) NOT NULL,
  `name` varchar(50) NOT NULL,
  `cardset_pid` int(10) NOT NULL,
  `cardtype` varchar(50) NOT NULL,
  `rarity` varchar(20) NOT NULL,
  `cost` int(3) NOT NULL,
  `attack` int(10) NOT NULL,
  `health` int(10) NOT NULL,
  `artist` varchar(50) NOT NULL,
  `collectible` tinyint(1) NOT NULL,
  `class_pid` int(10) NOT NULL,
  `race` varchar(50) NOT NULL,
  `durability` int(10) NOT NULL,
  `armor` int(10) NOT NULL,
  `multiclassgroup` varchar(50) NOT NULL,
  `text` text NOT NULL
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4;

ALTER TABLE `card`
  ADD PRIMARY KEY (`pid`);

ALTER TABLE `card`
  MODIFY `pid` int(10) NOT NULL AUTO_INCREMENT, AUTO_INCREMENT=1;
COMMIT;

【问题讨论】:

    标签: mysql sql json python-3.x utf-8


    【解决方案1】:

    \xa0 是空间的变体。它来自 Word 吗?

    但是,更相关的是,它不是 utf8;它是latin1 或其他非utf8 编码。您需要回到它的来源并将 that 更改为 utf8。

    或者,如果您的下一步只是将其放入 MySQL 表中,那么您需要说出有关客户端的真相——即它是以 latin1(而不是 utf8)编码的。完成此操作后,MySQL 将在 INSERT 期间为您处理转换。

    【讨论】:

    • 我在此处从“cards.collectible.json”文件中获得了 JSON 条目:api.hearthstonejson.com/v1/51510/enUS 当我在记事本或记事本++ 中打开下载的 json 文件时,它在屏幕的右下角显示文件编码是 utf-8,所以我绝对不确定 latin1 字符来自哪里(除非记事本和记事本 ++ 都以某种方式错误地检测到编码)自从提出我的问题以来,我一直在到处玩不同的功能和目前正在将文件作为二进制文件处理,这似乎适用于前 2 个条目
    • 但由于某种原因不是原始文件,所以我没有下载文件,而是在浏览器中打开它并使用记事本将所有字符复制到一个空的 txt 文件中,将该文件另存为 .json 并有效。在我用粘贴的字符加载新的 json 文件后,我的脚本生成了所需的输出。因此,即使它现在有效,我也不知道为什么。无论如何感谢瑞克的帮助。我还修改了我的脚本以使用 str().replace 转义 '\n',现在所有内容都正确插入到我的表中。
    猜你喜欢
    • 2011-05-06
    • 2020-08-08
    • 2020-08-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-04-23
    • 2010-10-30
    • 1970-01-01
    相关资源
    最近更新 更多