【问题标题】:Converting a string representation of a json/dict to something usable with python request将 json/dict 的字符串表示形式转换为可用于 python 请求的内容
【发布时间】:2018-03-19 00:35:31
【问题描述】:

我最终不得不放弃并寻求帮助。我正在检索具有 json 格式类型(但格式不正确 - 没有双引号)的文档(带有请求)并尝试将数据提取为普通字典。这就是我所拥有的:这很有效,并且会为您提供我试图从中提取数据的输出。

def test():
    url = "http://www.sgx.com/JsonRead/JsonstData"
    payload = {}
    payload['qryId'] = 'RSTIc'
    payload['timeout'] = 60
    header = {'User-Agent' : 'Mozilla/5.0 (compatible; MSIE 10.0; Linux i686; Trident/2.0)', 'Content-Type': 'text/html; charset=utf-8'}
    req = requests.get(url, headers = header, params = payload)
    print(req.url)
    prelim = req.content.decode('utf-8')
    print(type(prelim))
    print(prelim)

test()

在那之后我想要的是:(假设一个正常运行的字典)

for stock in prelim['items']:
    print(stock['N'])

这应该给我所有股票名称的列表。

我已经尝试过大多数 json 函数:prelim.json()、loads.、load.、dump.、dumps.、parse。似乎没有任何工作,因为数据格式不正确。我也试过 ast.literal_eval() 没有成功。我在 Stack Overflow 上尝试了一些示例,以将该字符串转换为正确的字典,但没有运气。我似乎无法转换该字符串以使其表现为正确的字典。如果你能指出我正确的方向,那将不胜感激。

好心人要求提供数据示例。来自上述请求的数据有点长,但我删除了一些“项目”,以便人们可以看到检索到的数据的一般外观。

{}&& {identifier:'ID', label:'As at 19-03-2018 8:38 AM',items:[{ID:0,N:'AscendasReit',SIP:'',NC: 'A17U',R:'',I:'',M:'',LT:0,C:0,VL:97.600,BV:485.300,B:'2.670',S:'2.670',SV:1009.100 ,O:0,H:0,L:0,V:259811.200,SC:'9',PV:2.660,P:0,P_:'X',V_:''}, {ID:1,N:'CapitaComTrust',SIP:'',NC:'C61U',R:'',I:'',M:'',LT:0,C:0,VL:126.349,BV :1467.300,B:'1.800',S:'1.800',SV:620.900,O:0,H:0,L:0,V:228691.690,SC:'9',PV:1.810,P:0,P_ :'X',V_:''}, {ID:2,N:'CapitaLand',SIP:'',NC:'C31',R:'',I:'',M:'',LT:0,C:0,VL:78.000,BV :184.900,B:'3.670',S:'3.670',SV:372.900,O:0,H:0,L:0,V:286026.000,SC:'9',PV:3.660,P:0,P_ :'X',V_:''}, {ID:28,N:'Wilmar Intl',SIP:'',NC:'F34',R:'CD',I:'',M:'',LT:0,C:0,VL:0.000 ,BV:32.000,B:'3.210',S:'3.210',SV:73.100,O:0,H:0,L:0,V:0.000,SC:'2',PV:3.220,P:0 ,P_:'',V_:''}, {ID:29,N:'YZJ Shipbldg SGD',SIP:'',NC:'BS6',R:'',I:'',M:'',LT:0,C:0,VL:0.000 ,BV:349.500,B:'1.330',S:'1.330',SV:417.700,O:0,H:0,L:0,V:0.000,SC:'2',PV:1.340,P:0 ,P_:'',V_:''}]}

根据最近的评论,我知道我可以这样做:

def test2():
    my_text = "{}&& {identifier:'ID', label:'As at 19-03-2018 8:38 AM',items:[{ID:0,N:'AscendasReit',SIP:'',NC:'A17U',R:'',I:'',M:'',LT:0,C:0,VL:97.600,BV:485.300,B:'2.670',S:'2.670',SV:1009.100,O:0,H:0,L:0,V:259811.200,SC:'9',PV:2.660,P:0,P_:'X',V_:''}, {ID:1,N:'CapitaComTrust',SIP:'',NC:'C61U',R:'',I:'',M:'',LT:0,C:0,VL:126.349,BV:1467.300,B:'1.800',S:'1.800',SV:620.900,O:0,H:0,L:0,V:228691.690,SC:'9',PV:1.810,P:0,P_:'X',V_:''}, {ID:2,N:'CapitaLand',SIP:'',NC:'C31',R:'',I:'',M:'',LT:0,C:0,VL:78.000,BV:184.900,B:'3.670',S:'3.670',SV:372.900,O:0,H:0,L:0,V:286026.000,SC:'9',PV:3.660,P:0,P_:'X',V_:''}, {ID:28,N:'Wilmar Intl',SIP:'',NC:'F34',R:'CD',I:'',M:'',LT:0,C:0,VL:0.000,BV:32.000,B:'3.210',S:'3.210',SV:73.100,O:0,H:0,L:0,V:0.000,SC:'2',PV:3.220,P:0,P_:'',V_:''}, {ID:29,N:'YZJ Shipbldg SGD',SIP:'',NC:'BS6',R:'',I:'',M:'',LT:0,C:0,VL:0.000,BV:349.500,B:'1.330',S:'1.330',SV:417.700,O:0,H:0,L:0,V:0.000,SC:'2',PV:1.340,P:0,P_:'',V_:''}]}"
    prelim = my_text.split("items:[")[1].replace("}]}", "}")
    temp_list = prelim.split(", ")
    end_list = []
    main_dict = {}
    for tok1 in temp_list:
        temp_dict = {}
        temp = tok1.replace("{","").replace("}","").split(",")
        for tok2 in temp:            
            my_key = tok2.split(":")[0]
            my_value = tok2.split(":")[1].replace("'","")
            temp_dict[my_key] = my_value
        end_list.append(temp_dict)    
    main_dict['items'] = end_list
    for stock in main_dict['items']:
        print(stock['N'])

test2()

这是期望的结果。我只是在问,是否有更简单(更优雅/pythonic)的方式来做到这一点。

【问题讨论】:

  • 请提供您的数据示例,特别是“格式错误”的细分。
  • 您应该显示数据的样子。这比你如何得到它更重要。检索数据的代码与实际问题完全无关。
  • o/p from sgx.com/JsonRead/JsonstData 不是 JSON 格式。它在字符串中。所以你需要手动转换成json可解析文本并做json.loads(converted_text)
  • 检查我的答案。让我知道它是否有效。

标签: python json dictionary python-requests


【解决方案1】:

您需要先将字符串转换为 JSON 可转换文本,然后使用json.loads 获取字典

prelim 不是 JSON 格式,值没有被 " 包围

  • 删除'{}&& '
  • 用"包围属性
  • 申请json.loads(new_text)获取字典表示

即

import requests, json
#replace tuples 
reps = (('identifier:', '"identifier":'),
        ('label:', '"label":'),
        ('items:', '"items":'),
        ('NC:', '"NC":'),
        ('ID:', '"ID":'),
        ('N:', '"N":'),
        ('SIP:', '"SIP":'),
        ('SC:', '"SC":'),
        ('R:', '"R":'),
        ('I:', '"I":'),
        ('M:', '"M":'),
        ('LT:', '"LT":'),
        ('C:', '"C":'),
        ('VL:', '"VL":'),
        ('BV:', '"BV":'),
        ('BL:', '"BL":'),
        ('B:', '"B":'),
        ('S:', '"S":'),
        ('SV:', '"SV":'),
        ('O:', '"O":'),
        ('H:', '"H":'),
        ('L:', '"L":'),
        ('PV:', '"PV":'),
        ('V:', '"V":'),
        ('P_:', '"P_":'),
        ('P:', '"P":'),
        ('V_:', '"V_":'))

#getting rid of invalid json text
prelim = prelim.replace('{}&& ', '')

#replacing single quotes with double quotes
prelim = prelim.replace("'", "\"")

print(prelim)
#reduce to get all replacements
dict_text = fn.reduce(lambda a, kv: a.replace(*kv), reps, prelim)
dic = json.loads(dict_text)
print(dic)

获取物品:

for x in dic['items']:
    print(x['N'])

输出:

2ndChance W200123
3Cnergy
3Cnergy W200528
800 Super
8Telecom^
A-Smart
A-Sonic Aero^
AA
....

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-01-19
    • 2011-05-30
    • 2021-02-20
    • 2015-10-27
    • 2019-03-27
    相关资源
    最近更新 更多