【问题标题】:Flattening a list of JSON strings with nested dict使用嵌套字典展平 JSON 字符串列表
【发布时间】:2020-12-24 17:53:57
【问题描述】:

我想转换以下list 的tuples:

[('1599324732926-0',
     {'data': '{"timestamp":1599324732.767,
                "receipt_timestamp":1599324732.9256856,
                "delta":true,
                "bid":{"338.9":0.06482,"338.67":3.95535},
                "ask":{"339.12":2.47578,"339.13":6.43172}
               }'
     }
 )
 ('1599324732926-1',
     {'data': '{"timestamp":1599324832.767,
                "receipt_timestamp":1599324832.9256856,
                "delta":true,
                "bid":{"338.8":0.06482,"338.57":3.95535},
                "ask":{"340.12":2.47578,"340.13":6.43172}
               }'
     }
 )
]

进入dicts 中的list 或Dataframe(任何一个,从一个到另一个并不复杂):

[{
  'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.9,
  'size': 0.06482},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.67,
  'size': 3.95535},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.66,
  'size': 16.78636},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.63,
  'size': 2.5},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.45,
  'size': 6.06071},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.38,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.95,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 338.96,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'ask',
  'price': 339.11,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 339.12,
  'size': 2.47578},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 339.13,
  'size': 6.43172},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 339.36,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 339.52,
  'size': 6.5},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 341.18,
  'size': 0.0},
 {'timestamp': 1599324732.767,
  'receipt_timestamp': 1599324732.9256856,
  'delta': True,
  'side': 'bid',
  'price': 341.19,
  'size': 0.0},
  ...
]

基本上,

  • 第一个 id 被删除(实际上,它保存在一个单独的列表中)。
  • data 中的数据是具有嵌套字典的 JSON 对象。
  • 诀窍在于“bid”和“ask”成为结果字典中名为“side”的键的值。
  • 嵌套字典“bid”和“ask”的键成为结果字典中名为“price”的键的值。
  • 一个名为“size”的键的价格保持值。

我能够分别处理列表中的每个 JSON 元素。 但是列表最多可以包含 600k 个元素。 我询问是否可以使用一些 pandas 或 numpy 函数将列表作为一个整体来处理以提高速度?

我查看了 pandas json_normalize(),但根据给出的示例,dict 的键是系统列,而在这种情况下,“价格”键成为“价格”列的值。

你知道我该怎么做吗?有什么方法可以对 JSON 列表进行第一次预处理,以便使用json_normalize() 对其进行进一步处理。

仅供参考,这是我可以编写的代码来分别处理列表中的每个元素,但我认为这不是正确的方向。下一步是将其封装在一个 for 循环中,与管理整个列表的解决方案相比,这将慢得多。

import json

data_light = ('1599324732926-0',
     {'data': '{"timestamp":1599324732.767, \
                "receipt_timestamp":1599324732.9256856,\
                "delta":true, \
                "bid":{"338.9":0.06482,"338.67":3.95535}, \
                "ask":{"339.12":2.47578,"339.13":6.43172} \
               }'
     }
 )

var=json.loads(data_light[1]['data'])
var_bid=var['bid']
var_ask=var['ask']
mylist=list(var_bid.items())+list(var_ask.items())

it = ['ask'] * len(var_ask) + ['bid'] * len(var_bid)

timestamp=var['timestamp']
receipt_timestamp=var['receipt_timestamp']
delta=var['delta']
midx = pd.MultiIndex.from_product([[timestamp], [receipt_timestamp], [delta],it], names=['timestamp', 'receipt_timestamp', 'delta', 'side'])

df=pd.DataFrame(mylist, index=midx, columns=['price', 'size'], dtype=float)
my_dict=df.reset_index().to_dict('records')

【问题讨论】:

    标签: python json pandas dictionary


    【解决方案1】:

    这不完全是您问题的答案,因为它不是 pandas 或 numpy 的实现,但我认为它应该可以满足您的需要。

    试试看multiprocessing.pool.Pool.map

    假设您有一个从原始列表接收元组并返回您想要的数据字典的函数。假设它的签名是这样的:

    def tuple_to_dict(input):
        # conversion code goes here
        return result_dict
    

    然后您可以像这样使用 multiprocessing.Pool():

    import multiprocessing
    
    
    if __name__ == '__main__':
    
        input_list = [...] # your input list
    
        with multiprocessing.Pool() as pool:
            result_list = pool.map(tuple_to_dict, input_list)
            print(result_list)
    

    注意:

    1. Pool() 对象的创建应该放在if __name__ == "__main__" 块或从那里调用的函数(递归) - 否则你会得到一个 RuntimeError

    2. with ... as... 放置在那里,以便在使用结束或失败时关闭 Pool 对象。如果您不使用“with / as”语法,请在 try/catch 块内使用它,并在其 finally 块中添加 pool.close() 语句以确保池已关闭。

    【讨论】:

    • 嗨,迈克尔,感谢您的提示。不幸的是,我的硬件已经到了极限,我不确定多处理是否真的可行。我主要是在寻找更快的代码。还是谢谢!最佳,
    【解决方案2】:
    • 迭代提取信息比使用pandas.json_normalize 更容易。
    • 如样本数据所示,data 的值为str 类型,必须转换为dict。
    • 主要任务是从'bid'和'ask'中提取每个keyvalue对,以创建单独的记录。
      • list-comprehension 执行创建单独记录的任务。
    import json
    import pandas
    
    # list of tuples, where the value of data, is a string
    transaction_data = [('1599324732926-0', {'data': '{"timestamp":1599324732.767, "receipt_timestamp":1599324732.9256856, "delta":true, "bid":{"338.9":0.06482,"338.67":3.95535}, "ask":{"339.12":2.47578,"339.13":6.43172}}'}),
                        ('1599324732926-1', {'data': '{"timestamp":1599324732.767, "receipt_timestamp":1599324732.9256856, "delta":true, "bid":{"338.9":0.06482,"338.67":3.95535}, "ask":{"339.12":2.47578,"339.13":6.43172}}'}),
                        ('1599324732926-2', {'data': '{"timestamp":1599324732.767, "receipt_timestamp":1599324732.9256856, "delta":true, "bid":{"338.9":0.06482,"338.67":3.95535}, "ask":{"339.12":2.47578,"339.13":6.43172}}'})]
    
    # create a list of lists for each transaction data
    # split each side, key value pair into a separate list
    data_key_list = [['timestamp', 'receipt_timestamp', 'delta', 'side', 'price', 'size']]
    
    for v in transaction_data:  # # iterate through each transaction
        data = json.loads(v[1]['data'])  # convert the string to a dict
        for side in ['bid', 'ask']:  # extract each key, value pair as a separate record
            data_key_list += [[data['timestamp'], data['receipt_timestamp'], data['delta'], side, float(k), v] for k, v in data[side].items()]
    
    # create a dataframe
    df = pd.DataFrame(data_key_list[1:], columns=data_key_list[0])
    
    # display(df.head())
         timestamp  receipt_timestamp  delta side   price     size
    0  1.59932e+09        1.59932e+09   True  bid   338.9  0.06482
    1  1.59932e+09        1.59932e+09   True  bid  338.67  3.95535
    2  1.59932e+09        1.59932e+09   True  ask  339.12  2.47578
    3  1.59932e+09        1.59932e+09   True  ask  339.13  6.43172
    4  1.59932e+09        1.59932e+09   True  bid   338.9  0.06482
    

    转换为字典列表

    df.to_dict(orient='records')
    
    [out]:
    [{'timestamp': 1599324732.767,
      'receipt_timestamp': 1599324732.9256856,
      'delta': True,
      'side': 'bid',
      'price': 338.9,
      'size': 0.06482},
     {'timestamp': 1599324732.767,
      'receipt_timestamp': 1599324732.9256856,
      'delta': True,
      'side': 'bid',
      'price': 338.67,
      'size': 3.95535},
     {'timestamp': 1599324732.767,
      'receipt_timestamp': 1599324732.9256856,
      'delta': True,
      'side': 'ask',
      'price': 339.12,
      'size': 2.47578},
     {'timestamp': 1599324732.767,
      'receipt_timestamp': 1599324732.9256856,
      'delta': True,
      'side': 'ask',
      'price': 339.13,
      'size': 6.43172},
     ...]
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-07-12
      • 1970-01-01
      • 2021-11-05
      • 1970-01-01
      • 2014-02-16
      • 2019-05-03
      • 2018-10-30
      • 1970-01-01
      相关资源
      最近更新 更多