【发布时间】:2017-01-30 15:07:53
【问题描述】:
我在几个目录中有 json 数据文件,我想导入 Pandas 进行一些数据分析。 json 的格式取决于目录名称中定义的类型。例如,
dir1_typeA/
file1
file2
...
dir1_typeB/
file1
file2
...
dir2_typeB/
file1
...
dir2_typeA/
file1
file2
每个file 都包含一个复杂的嵌套 json 字符串,它将是 DataFrame 的一行。我将为每个 TypeA 和 TypeB 提供两个数据框。稍后我会在需要时附加它们。
所以,到目前为止,我已经获得了 os.walk 所需的所有文件路径,并且正在尝试通过
import os
from glob import glob
PATH = 'dir/filepath'
files = [y for x in os.walk(PATH) for y in glob(os.path.join(x[0], 'file*'))]
for file in files:
with open(issuefile, 'r') as f:
data = f.read()
data_json = json_normalize(json.loads(data))
type = ' '.join(issuefile.split('/')[3]
data_json['type'] = type
# append to data frame for typeA and typeB
if 'typeA' in type:
# append to typeA dataframe
else:
# append to typeB dataframe
还有一个额外的问题,即目录中的文件可能具有稍微不同的字段。例如,file1 在dir1_typeA 中可能有更多的字段file2。因此,我还需要在每种类型的数据框中适应这种动态特性。
如何创建这两个数据框?
【问题讨论】:
标签: python json pandas dataframe