【问题标题】:JSON normalize with pandas - list indices need to be int使用 pandas 规范化 JSON - 列表索引需要是 int
【发布时间】:2018-03-27 01:42:24
【问题描述】:

我正在处理一个大型 JSON,我想将其转换为 csv 以进行进一步分析。 当我使用 json_normalize 构建表时,出现以下错误:

Traceback(最近一次调用最后一次):

文件“/Users/Home/Downloads/JSONtoCSV/easybill.py”,第 30 行,在 “status”、“text”、“text_prefix”、“title”、“type”、“use_shipping_address”、“vat_option”

文件“/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/site-packages/pandas/io/json/normalize.py”,第 248 行,在 json_normalize _recursive_extract(data, record_path, {}, level=0)

文件“/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/site-packages/pandas/io/json/normalize.py”,第 235 行,在 _recursive_extract meta_val = _pull_field(obj, val[level:])

文件“/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/site-packages/pandas/io/json/normalize.py”,第 169 行,在 _pull_field 结果 = 结果[字段]

TypeError: 列表索引必须是整数,而不是 str

在第一步中,我使用更小/更精简的 JSON 进行了许多测试以进行代码验证。现在,当我为完整的 JSON 组装所有内容时,我收到了这条错误消息。

我该如何解决这个问题?我正在尝试使用 pandas 实现标准化,如下所示:http://pandas.pydata.org/pandas-docs/stable/io.html#normalization

这是我到目前为止的代码。感谢您的帮助!

编辑:这是 JSON 源:https://pastebin.com/muGBPWv8

# -*- coding: utf-8 -*-
import pandas
import json
import sys
reload(sys)
sys.setdefaultencoding('utf-8')
from pandas.io.json import json_normalize

# Paths
json_file_path = "/Users/Home/Downloads/JSONtoCSV/JSON-Files/Seite0.json"
csv_file_path = "/Users/Home/Downloads/JSONtoCSV/CSV-files/Seite0.csv"
node = "items"

# JSON file open, no pagination information
with open(json_file_path) as f:
    rawjson = json.load(f)
data = rawjson[node]

# remove "number" because it causes errors in pandas.
good_data = eval(repr(data).replace("number", "numbr"))

# normalization
norm_data =  json_normalize(good_data, "items", [
["address","city"], ["address","company_name"], ["address","country"], ["address","first_name"], ["address","last_name"], ["address","personal"], ["address","salutation"], ["address","street"], ["address","suffix_1"], ["address","suffix_2"], ["address","title"], ["address","zip_code"], 
"amount", "amount_net", "attachment_ids", "bank_debit_form", "cancel_id", "cash_allowance", "cash_allowance_days", "cash_allowance_text", "contact_id", "contact_label", "contact_text", "created_at", "currency", "customer_id", "discount", "discount_type", "document_date", "due_date", "edited_at", "external_id", "grace_period", "id", "is_archive", "is_draft", "is_replica",
["items","booking_account"], ["items","cost_price_charge"], ["items","cost_price_charge_type"], ["items","cost_price_net"], ["items","cost_price_total"], ["items","description"], ["items","discount"], ["items","discount_type"], ["items","export_cost_1"], ["items","export_cost_2"], ["items","id"], ["items","numbr"], ["items","position"], ["items","position_id"], ["items","quantity"], ["items","quantity_str"], ["items","serial_number"], ["items","serial_number_id"], ["items","single_price_gross"], ["items","single_price_net"], ["items","total_price_gross"], ["items","total_price_net"], ["items","total_vat"], ["items","type"], ["items","unit"], ["items","vat_percent"],
"label_address", "label_address", "login_id", "numbr", "paid_amount", "paid_at", "pdf_pages", "pdf_template", "project_id", "ref_id", "replica_url",
["service_date","type"], ["service_date","date"], ["service_date","date_from"], ["service_date","date_to"], ["service_date","text"], 
"status", "text", "text_prefix", "title", "type", "use_shipping_address", "vat_option"
])

# save to csv
norm_data.to_csv(csv_file_path, sep=";")

【问题讨论】:

    标签: python json pandas


    【解决方案1】:

    我发现您的代码存在几个问题:

    1. 您的元数据 ID 存在冲突。例如,您将 'id' 作为元数据(级别 1 项),并将 'id' 作为 'items' 的一个元素。这可以通过给json_normalize 提供第三个参数来解决,比如

      json_normalize(good_data, "items", [...], "meta."

    2. json_normalize 期望元数据存储在字典中(可能是字典中的递归),但您有值类型为 list 的项目,例如 attachment_ids。目前json_normalize似乎无法处理它们。

    3. 另外,json_normalize 似乎无法处理空字典,例如 "label_address": {}。

    4. 1234563 987654334@)。

    考虑到json_normalize 的问题,我不想用它来解决你的问题,而只是写下简单的命令式代码(带有循环/列表推导),从你的 JSON 中创建一个表(列表列表)然后从该表中创建 pandas 数据框。

    【讨论】:

    • 您好 - 感谢您的评论!我使用json_normalize 只是因为我需要使用平面表进行标准化,结果如下所示:pandas.pydata.org/pandas-docs/stable/io.html#normalization“Rick Scott”多次出现在决赛桌中 - 但在源 json 中只有一次。你对我有什么建议吗?我对 python 完全陌生 - 所以要通过 :(
    猜你喜欢
    • 2020-12-11
    • 1970-01-01
    • 2018-05-09
    • 2021-07-11
    • 2021-04-26
    • 1970-01-01
    • 1970-01-01
    • 2023-03-23
    • 2020-12-15
    相关资源
    最近更新 更多