【问题标题】:Python: Reading 200k JSON files into a Pandas DataframePython:将 200k JSON 文件读入 Pandas 数据框
【发布时间】:2017-09-25 10:10:53
【问题描述】:

我是 非常 python 新手( 1000 变化,但那 50 是 1000 的子集。

这是一个 json 文件的 sn-p:

{
"study_type" : "Observational",
"intervention.intervention_type" : "Device",
"primary_outcome.time_frame" : "24 months",
"primary_completion_date.type" : "Actual",
"design_info.primary_purpose" : "Diagnostic",
"design_info.secondary_purpose" : "Intervention",
"start_date" : "January 2014",
"end_date" : "March 2014",
"overall_status" : "Completed",
"location_countries.country" : "United States",
"location.facility.name" : "Generic Institution",
}

我们的目标是获取这些 JSON 文件的主数据库,清理各个列,对这些列运行描述性统计数据,然后创建一个最终的清理数据库。

我来自 SAS 背景,所以我的想法是使用 pandas 并创建一个(非常)大的数据框。过去一周我一直在梳理堆栈溢出问题,并从中吸取了一些教训,但我觉得必须有一种方法可以让这种方式更有效率。

以下是我到目前为止编写的代码 - 它运行,但非常慢(我估计即使在消除了以“结果”开头的不需要的输入属性/列之后,也需要数天甚至数周才能运行)。

此外,我将字典转换为最终表格的尴尬方式是在列名上方留下列索引号,我无法弄清楚如何删除。

import json, os
import pandas as pd    
from copy import deepcopy

path_to_json = '/home/ubuntu/json_flat/'

#Gets list of files in directory with *.json suffix
list_files = [pos_json for pos_json in os.listdir(path_to_json) if pos_json.endswith('.json')]

#Initialize series
df_list = []

#For every json file found
for js in list_files:

    with open(os.path.join(path_to_json, js)) as data_file:
        data = json.loads(data_file.read())                         #Loads Json file into dictionary
        data_file.close()                                           #Close data file / remove from memory

        data_copy = deepcopy(data)                                  #Copies dictionary file
        for k in data_copy.keys():                                  #Iterate over copied dictionary file
            if k.startswith('result'):                              #If field starts with "X" then delete from dictionary
                del data[k]
        df = pd.Series(data)                                        #Convert Dictionary to Series
        df_list.append(df)                                          #Append to empty series  
        database = pd.concat(df_list, axis=1).reset_index()         #Concatenate series into database

output_db = database.transpose()                                    #Transpose rows/columns
output_db.to_csv('/home/ubuntu/output/output_db.csv', mode = 'w', index=False)

非常感谢任何想法和建议。如果它更有效并且仍然允许我们实现上述目标,我完全愿意完全(在 python 中)使用不同的技术或方法。

谢谢!

【问题讨论】:

  • 请注意,您可以使用json.load() (docs) 直接读取文件。无需添加read() 并执行json.loads()。
  • 另外,你为什么不把所有的 json 读入一个大字典,然后将整个字典转换成一个可以写入文件的 pandas DataFrame(参见 here)。跨度>
  • 谢谢帕特里克!欣赏小费,我会做出改变。它似乎对运行时间影响不大,但每一种效率都有帮助。
  • @patrick 刚刚在我回复时看到了您的第二个帖子。让我试一试。

标签: python json pandas dataframe


【解决方案1】:

您最关键的性能错误可能是这样的:

database = pd.concat(df_list, axis=1).reset_index()

您在循环中执行此操作,每次向df_list 添加一件事,然后再次连接。但是直到最后都没有使用这个“数据库”变量,所以你可以在循环外只做一次这一步。

对于 Pandas,循环中的“concat”是一个巨大的反模式。在循环中构建您的列表,连接一次。

第二件事是你也应该使用 Pandas 来读取 JSON 文件:http://pandas.pydata.org/pandas-docs/stable/generated/pandas.read_json.html

保持简单。编写一个获取路径、调用pd.read_json()、删除不需要的行(series.str.startswith())等的函数。

一旦你的工作正常,下一步就是检查你是否受到 CPU 限制(CPU 使用率 100%)或 I/O 限制(CPU 使用率远低于 100%)。

【讨论】:

  • 感谢约翰,仅一行移动就显着改变了运行时间。对于我的一生,我无法让startswith()函数与这个系列一起工作(我可以在这些文件上使用“read_json()”的唯一方法是:typ ='series'和orient ='records'w /o 错误)。 -------------------------------------------------- --------------- if not data.str.startswith('results'):(返回 ValueError)
【解决方案2】:

我尝试以更简洁的方式复制您的方法,减少复制和附加。它适用于您提供的示例数据,但不知道您的数据集中是否还有其他复杂性。你可以试试这个,我希望cmets帮助。

import json
import os
import pandas
import io


path_to_json = "XXX"

list_files = [pos_json for pos_json in os.listdir(path_to_json) if pos_json.endswith('.json')]

#set up an empty dictionary
resultdict = {}

for fili in list_files:
    #the with avoids the extra step of closing the file
    with open(os.path.join(path_to_json, fili), "r") as inputjson:
        #the dictionary key is set to filename here, but you could also use e.g. a counter
        resultdict[fili] = json.load(inputjson)
        """
        you can exclude stuff here or later via dictionary comprehensions: 
        http://stackoverflow.com/questions/1747817/create-a-dictionary-with-list-comprehension-in-python
        e.g. as in your example code
        resultdict[fili] = {k:v for k,v in json.load(inputjson).items() if not k.startswith("result")}
        """

#put the whole thing into the DataFrame     
dataframe = pandas.DataFrame(resultdict)

#write out, transpose for desired format
with open("output.csv", "w") as csvout:
    dataframe.T.to_csv(csvout)

【讨论】:

  • 帕特里克,这很好用!它很好地摆脱了列索引值。我遇到问题(JSONDecodeError)的地方是我取消注释排除标准。看起来生成的resultdict[fili] 在转换为数据帧之前有一个 json 文件名的键和包含键/值对的值。再次感谢您的任何想法/建议。
  • @RDara 这样做时,您需要注释掉没有排除的第一步。做到了吗?
  • 抱歉,您指的是哪一步?基本上我在以下行中添加:resultdict[file] = {k: v for k,v in json.load(inputjson).items() if not k.startswith("location")}(更改为“位置”,因为它在上面的示例 JSON 文件中).. 感谢您的耐心等待!
  • 补充一点,最初花费我 90 多分钟的时间在 10k 个文件的子集上减少到几秒钟。刚刚在完整的 200k+ JSON 文件(提取字段子集)上进行了测试,这似乎也运行得相当快。真是太感谢你了!!
  • 您可以将文件名传递给to_csv() - 不需要open()。
猜你喜欢
  • 2020-08-27
  • 2019-11-25
  • 2021-09-07
  • 2021-01-21
  • 1970-01-01
  • 1970-01-01
  • 2018-03-23
  • 2019-10-13
  • 2019-07-02
相关资源
最近更新 更多