【问题标题】:How to import normalised json data from several files into a pandas dataframe?如何将多个文件中的规范化 json 数据导入 pandas 数据框?
【发布时间】:2017-01-30 15:07:53
【问题描述】:

我在几个目录中有 json 数据文件,我想导入 Pandas 进行一些数据分析。 json 的格式取决于目录名称中定义的类型。例如,

dir1_typeA/
  file1
  file2
  ...
dir1_typeB/
  file1
  file2
  ...
dir2_typeB/
  file1
  ...
dir2_typeA/
  file1
  file2

每个file 都包含一个复杂的嵌套 json 字符串,它将是 DataFrame 的一行。我将为每个 TypeA 和 TypeB 提供两个数据框。稍后我会在需要时附加它们。

所以,到目前为止,我已经获得了 os.walk 所需的所有文件路径,并且正在尝试通过

    import os
    from glob import glob

    PATH = 'dir/filepath'
    files = [y for x in os.walk(PATH) for y in glob(os.path.join(x[0], 'file*'))]

    for file in files:
        with open(issuefile, 'r') as f:
            data = f.read()

        data_json = json_normalize(json.loads(data))
        type = ' '.join(issuefile.split('/')[3]
        data_json['type'] = type
        # append to data frame for typeA and typeB
        if 'typeA' in type:
            # append to typeA dataframe
        else:
            # append to typeB dataframe

还有一个额外的问题,即目录中的文件可能具有稍微不同的字段。例如,file1 在dir1_typeA 中可能有更多的字段file2。因此,我还需要在每种类型的数据框中适应这种动态特性。

如何创建这两个数据框?

【问题讨论】:

标签: python json pandas dataframe


【解决方案1】:

我认为您应该先将文件连接在一起,然后再将它们读入 pandas,这是您在 bash 中的操作方式(您也可以在 Python 中进行操作):

cat `find *typeA` > typeA
cat `find *typeB` > typeB

然后你可以使用io.json.json_normalize将其导入pandas:

import json
with open('typeA') as f:
    data = [json.loads(l) for l in f.readlines()]
    dfA = pd.io.json.json_normalize(data)

dfA

#          that this.first this.second
# 0  otherthing      thing       thing
# 1  otherthing      thing       thing
# 2  otherthing      thing       thing

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2019-11-25
    • 2020-12-15
    • 2019-10-13
    • 2019-09-01
    • 2012-09-13
    • 2021-04-18
    • 2021-01-21
    相关资源
    最近更新 更多