【问题标题】:Nested dictionary to pandas DataFrame -嵌套字典到 pandas DataFrame -
【发布时间】:2022-01-01 20:43:52
【问题描述】:

我的嵌套 JSON -

Instance1 [{'datapoints': [{'statistic': 'Minimum', 'timestamp': '2021-08-31 06:50:00.000000', 'value': 59.03}, {'statistic': '最小值','时间戳':'2021-08-18 02:50:00.000000','值':59.37},{'统计':'最小值','时间戳':'2021-08-24 16:50: 00.000000', 'value': 58.84},, 'metric': 'VolumeIdleTime', 'unit': 'Seconds'}]

Instance2 [{'datapoints': [{'statistic': 'Minimum', 'timestamp': '2021-08-31 06:50:00.000000', 'value': 60}, {'statistic': '最小值','时间戳':'2021-08-18 02:50:00.000000','值':55.45},{'统计':'最小值','时间戳':'2021-08-24 16:50: 00.000000','值':54.16},{'统计':'最小值','时间戳':'2021-08-06 07:50:00.000000','值':50.03},{'统计':'最小值','时间戳':'2021-08-04 22:50:00.000000','值':60},{'统计':'最小值','时间戳':'2021-08-26 01:50:00.000000 ', 'value': 60.34}, 'metric': 'VolumeIdleTime', 'unit': 'Seconds'}]

Instance3 [{'datapoints': [{'statistic': 'Minimum', 'timestamp': '2021-08-31 06:50:00.000000', 'value': 60}, {'statistic': '最小值','时间戳':'2021-08-18 02:50:00.000000','值':38.12},{'统计':'最小值','时间戳':'2021-08-24 16:50: 00.000000','值':42.31},{'统计':'最小值','时间戳':'2021-08-06 07:50:00.000000','值':45.22},{'统计':'最小值','时间戳':'2021-08-04 22:50:00.000000','值':40.51},{'统计':'最小值','时间戳':'2021-08-26 01:50:00.000000 ', 'value': 34.35}, {'statistic': 'Minimum', 'timestamp': '2021-08-11 12:50:00.000000', 'value': 46.33},'metric': 'VolumeIdleTime', “单位”:“秒”}]

还有更多实例详细信息(接近 8K 实例信息)

我调用了以下函数来反规范化我的嵌套 JSON:

def flatten_nested_json_df(df): df = df.reset_index() # 重置索引 & 设置一个从 0 到数据长度的整数列表作为索引。 s = (df.applymap(type) == list).all() # 检查所有值是否为真 list_columns = s[s].index.tolist()

s = (df.applymap(type) == dict).all()
dict_columns = s[s].index.tolist()


while len(list_columns) > 0 or len(dict_columns) > 0:
    new_columns = []

    for col in dict_columns:
        horiz_exploded = pd.json_normalize(df[col]).add_prefix(f'{col}.')
        horiz_exploded.index = df.index
        df = pd.concat([df, horiz_exploded], axis=1).drop(columns=[col])
        new_columns.extend(horiz_exploded.columns) # inplace

    for col in list_columns:
        #print(f"exploding: {col}")
        df = df.drop(columns=[col]).join(df[col].explode().to_frame())
        new_columns.append(col)

    s = (df[new_columns].applymap(type) == list).all()
    list_columns = s[s].index.tolist()

    s = (df[new_columns].applymap(type) == dict).all()
    dict_columns = s[s].index.tolist()
return df

然后是 flatten_nested_json_df(df)

我的数据框的前 10 行编码工作正常。 eg- Less_metric= met_data_aug.iloc[:10,:] flatten_nested_json_df(Less_metric) # 解析在所有层工作并提取所有列

输出: _id metrics.metric metrics.unit metrics.datapoints.statistic metrics.datapoints.timestamp metrics.datapoints.value 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-31 06:50:00.000000 59.03 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-18 02:50:00.000000 59.37 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-24 16:50:00.000000 58.84 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-06 07:50:00.000000 59.41 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-04 22:50:00.000000 59.37 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小值 2021-08-26 01:50:00.000000 57.21 0 612e09dd3d2001b69bd0dad1 VolumeIdleTime 秒最小 2021-08-11 12:50:00.000000 59.43

如果我选择 1000 行: 例如- med_metric=met_data_aug.iloc[:1000,:] 或者如果我选择随机线 ran_metric =met_data_aug.sample(10)

flatten_nested_json_df(med_metric) or flatten_nested_json_df(ran_metric) # 解析不全,没有提取所有列数据。

index   metrics.metric  metrics.unit    metrics.datapoints

0 13445 VolumeIdleTime 秒 {'statistic': 'Minimum', 'timestamp': '2021-08-07 18:40:00.000000', 'value': 59.98} 0 13445 VolumeIdleTime 秒 {'statistic': 'Minimum', 'timestamp': '2021-08-14 08:40:00.000000', 'value': 59.99} 0 13445 VolumeIdleTime 秒 {'statistic': 'Minimum', 'timestamp': '2021-08-13 06:40:00.000000', 'value': 59.99} 0 13445 VolumeIdleTime 秒 {'statistic': 'Minimum', 'timestamp': '2021-08-19 20:40:00.000000', 'value': 59.98} 0 13445 VolumeIdleTime Seconds {'statistic': 'Minimum', 'timestamp': '2021-08-26 10:40:00.000000', 'value': 59.98}

实例上的任何 NAN 值或列差异都会导致编码中断?

如何解决这个问题?- 请指导!!

【问题讨论】:

    标签: python pandas dictionary


    【解决方案1】:

    IIUC 你需要的是json_normalize。将datapoints设置为record_path和metric和unit设置为meta:

    data = [{'datapoints': [{'statistic': 'Minimum', 'timestamp': '2021-08-31 06:50:00.000000', 'value': 59.03},{'statistic': 'Minimum', 'timestamp': '2021-08-18 02:50:00.000000', 'value': 59.37}, {'statistic': 'Minimum', 'timestamp': '2021-08-24 16:50:00.000000', 'value': 58.84}],'metric': 'VolumeIdleTime', 'unit': 'Seconds'}]
    df = pd.json_normalize(data, record_path="datapoints", meta=["metric", "unit"])
    print(df)
    

    输出:

      statistic                   timestamp  value          metric     unit
    0   Minimum  2021-08-31 06:50:00.000000  59.03  VolumeIdleTime  Seconds
    1   Minimum  2021-08-18 02:50:00.000000  59.37  VolumeIdleTime  Seconds
    2   Minimum  2021-08-24 16:50:00.000000  58.84  VolumeIdleTime  Seconds
    

    【讨论】:

    • 您好 Tranbi,感谢您花时间解决我的问题。我尝试同样适用于我的嵌套 JSON 。但我收到以下错误“TypeError:列表索引必须是整数或切片,而不是 str”
    • 250 # GH 31507 GH 30145, GH 26284 如果结果未列出,则引发 TypeError 如果不是 ~\anaconda3\lib\site-packages\pandas\io\json_normalize.py in _pull_field(js, spec ) 237 result = result[field] 238 else: --> 239 result = result[spec] 240 return result 241 TypeError: list indices must be integers or slices, not str
    • 它与我在答案中使用的不同吗?你能在你的问题中发布你的 json 吗?
    • @SathishKumar 这怎么不能回答问题?请提供有关您面临的问题的详细信息(在问题中格式化,否则很难作为评论阅读)
    猜你喜欢
    • 2022-01-02
    • 2017-12-26
    • 2015-08-03
    • 2019-06-27
    • 1970-01-01
    • 1970-01-01
    • 2020-02-18
    • 2019-12-01
    • 1970-01-01
    相关资源
    最近更新 更多