【发布时间】:2020-11-19 04:05:52
【问题描述】:
我在 pandas 中加载一个大的 JSON 行文件时遇到问题,主要是因为我需要在使用 pd.read_json 后“展平”其中一个结果列 例如,对于这个 JSON 行:
{"user_history": [{"event_info": 248595, "event_timestamp": "2019-10-01T12:46:03.145-0400", "event_type": "view"}, {"event_info": 248595, "event_timestamp": "2019-10-01T13:21:50.697-0400", "event_type": "view"}], "item_bought": 1909110}
我需要像这样在熊猫中加载 2 行 4 列:
+--------------+--------------------------------+--------------+---------------+
| "event_info" | "event_timestamp" | "event_type" | "item_bought" |
+--------------+--------------------------------+--------------+---------------+
| 248595 | "2019-10-01T12:46:03.145-0400" | "view" | 1909110 |
| 248595 | "2019-10-01T13:21:50.697-0400" | "view" | 1909110 |
+--------------+--------------------------------+--------------+---------------+
问题是,考虑到文件的大小(413000+ 行,超过 1GB),我设法做到这一点的任何方法都不够快。我正在尝试一种相当基本的方法,遍历加载的数据帧,创建字典并将值附加到空数据帧:
history_df = pd.read_json('data/train_dataset.jl', lines=True)
history_df['index1'] = history_df.index
normalized_history = pd.DataFrame()
for index, row in history_df.iterrows():
for dic in row['user_history']:
dic['index1'] = row['index1']
dic['item_bought'] = row['item_bought']
normalized_history = normalized_history.append(dic, ignore_index=True)
所以问题是哪种方法最快?有什么办法不迭代整个 history_df 数据框?
提前谢谢你
【问题讨论】:
-
您能否提供您尝试但不太奏效的方法?另外,我们在这里讨论的文件大小是多少?
-
如果你能包含你目前拥有的代码,那将有助于审查和调试。此外,当您指出“最快的方式”时,这在很大程度上取决于您运行代码的环境。
-
@HubertGrzeskowiak 很抱歉错过了这个,用信息编辑了问题
-
@etch_45 已添加代码,感谢您的宝贵时间
标签: python json python-3.x pandas