【发布时间】:2017-09-25 10:10:53
【问题描述】:
我是 非常 python 新手( 1000 变化,但那 50 是 1000 的子集。
这是一个 json 文件的 sn-p:
{
"study_type" : "Observational",
"intervention.intervention_type" : "Device",
"primary_outcome.time_frame" : "24 months",
"primary_completion_date.type" : "Actual",
"design_info.primary_purpose" : "Diagnostic",
"design_info.secondary_purpose" : "Intervention",
"start_date" : "January 2014",
"end_date" : "March 2014",
"overall_status" : "Completed",
"location_countries.country" : "United States",
"location.facility.name" : "Generic Institution",
}
我们的目标是获取这些 JSON 文件的主数据库,清理各个列,对这些列运行描述性统计数据,然后创建一个最终的清理数据库。
我来自 SAS 背景,所以我的想法是使用 pandas 并创建一个(非常)大的数据框。过去一周我一直在梳理堆栈溢出问题,并从中吸取了一些教训,但我觉得必须有一种方法可以让这种方式更有效率。
以下是我到目前为止编写的代码 - 它运行,但非常慢(我估计即使在消除了以“结果”开头的不需要的输入属性/列之后,也需要数天甚至数周才能运行)。
此外,我将字典转换为最终表格的尴尬方式是在列名上方留下列索引号,我无法弄清楚如何删除。
import json, os
import pandas as pd
from copy import deepcopy
path_to_json = '/home/ubuntu/json_flat/'
#Gets list of files in directory with *.json suffix
list_files = [pos_json for pos_json in os.listdir(path_to_json) if pos_json.endswith('.json')]
#Initialize series
df_list = []
#For every json file found
for js in list_files:
with open(os.path.join(path_to_json, js)) as data_file:
data = json.loads(data_file.read()) #Loads Json file into dictionary
data_file.close() #Close data file / remove from memory
data_copy = deepcopy(data) #Copies dictionary file
for k in data_copy.keys(): #Iterate over copied dictionary file
if k.startswith('result'): #If field starts with "X" then delete from dictionary
del data[k]
df = pd.Series(data) #Convert Dictionary to Series
df_list.append(df) #Append to empty series
database = pd.concat(df_list, axis=1).reset_index() #Concatenate series into database
output_db = database.transpose() #Transpose rows/columns
output_db.to_csv('/home/ubuntu/output/output_db.csv', mode = 'w', index=False)
非常感谢任何想法和建议。如果它更有效并且仍然允许我们实现上述目标,我完全愿意完全(在 python 中)使用不同的技术或方法。
谢谢!
【问题讨论】:
标签: python json pandas dataframe