【问题标题】:Python 3 - Load JSON data into my .csv -filePython 3 - 将 JSON 数据加载到我的 .csv 文件中
【发布时间】:2018-06-16 06:04:41
【问题描述】:

我实际上尝试在 data.csv 文件中编写 JSON。我尝试了stackoverflow的以下解决方案: How do I write a Python dictionary to a csv file?

所以我想出了这些:

with open("data/dataGold.csv", 'w') as f:
    w = csv.DictWriter(f, ['data']['user']['repositories']['nodes'], extrasaction='ignore')
    w.writeheader()
    w.writerow(response)
    w.writerow([data['data']['user']['repositories']['nodes']['name'],
              data['data']['user']['repositories']['nodes']['forkCount'],
              data['data']['user']['repositories']['nodes']['issues']])

我的“dict”类型的响应变量是:

{'data': {'user': {'name': 'Markus Goldstein',
                   'repositories': {'nodes': [{'forkCount': 0,
                                               'issues': {'totalCount': 0},
                                               'name': 'repache'},
                                              {'forkCount': 4,
                                               'issues': {'totalCount': 3},
                                               'name': 'nf-hishape'},
                                              {'forkCount': 4,
                                               'issues': {'totalCount': 7},
                                               'name': 'ip-countryside'},
                                              {'forkCount': 42,
                                               'issues': {'totalCount': 29},
                                               'name': 'bonesi'},
                                              {'forkCount': 13,
                                               'issues': {'totalCount': 3},
                                               'name': 'rapidminer-anomalydetection'},
                                              {'forkCount': 0,
                                               'issues': {'totalCount': 0},
                                               'name': 'rapidminer-studio'}]}}}}

有一个 TypeError 表示不允许使用索引。我认为那是因为我使用了 ['data']['user']['repositories']['nodes']。

我发布上面链接的解决方案有效,因为没有嵌套的 Dict/JSON。所以我不知道在我的情况下如何使用嵌套的 Dict/JSON

所以我的目标是包含 name、forkCount 和 issues 作为标题的 CSV。下一行是不同 repo 的值。

愿有人可以帮助我,并为我的英语不好感到抱歉-.- 谢谢!

【问题讨论】:

    标签: json python-3.x csv dictionary


    【解决方案1】:

    下面给出的应该可以正常工作,

    1] 我编写的额外循环将您的结构更改为删除 issues 下的字典并将 total_counts 的值存储在 issues 下的结构> 这样 CSV 就干净了。

    2] 我在这里使用 deepcopy 是因为我不想修改原始数据结构,因此我没有使用引用,而是使用它的 deepcopy。

    3] 类型转换 wt_csv[0].keys() 到 list 因为.keys() 函数在 python 3 中返回 dict_keys 而不是列表

    import csv
    import json
    import copy
    
    i_dict = {'data': {'user': {'name': 'Markus Goldstein',
                               'repositories': {'nodes': [{'forkCount': 0,
                                                           'issues': {'totalCount': 0},
                                                           'name': 'repache'},
                                                          {'forkCount': 4,
                                                           'issues': {'totalCount': 3},
                                                           'name': 'nf-hishape'},
                                                          {'forkCount': 4,
                                                           'issues': {'totalCount': 7},
                                                           'name': 'ip-countryside'},
                                                          {'forkCount': 42,
                                                           'issues': {'totalCount': 29},
                                                           'name': 'bonesi'},
                                                          {'forkCount': 13,
                                                           'issues': {'totalCount': 3},
                                                           'name': 'rapidminer-anomalydetection'},
                                                          {'forkCount': 0,
                                                           'issues': {'totalCount': 0},
                                                           'name': 'rapidminer-studio'}]}}}}
    
    
    wt_csv = copy.deepcopy(i_dict['data']['user']['repositories']['nodes'])
    
    for wc in wt_csv:
        wc['issues'] = wc['issues']['totalCount']
    
    with open('dataGold.csv', 'w') as output_file:
        dict_writer = csv.DictWriter(output_file, fieldnames=list(wt_csv[0].keys()))
        dict_writer.writeheader()
        dict_writer.writerows(wt_csv)
    

    如果有不清楚的地方,请在 cmets 中告诉我。

    【讨论】:

    • 这实际上是完美的 :) 谢谢 :) 由于您的清晰描述,我都明白了。
    【解决方案2】:
    Markus = {'data': {'user': {'name': 'Markus Goldstein',
                       'repositories': {'nodes': [{'forkCount': 0,
                                                   'issues': {'totalCount': 0},
                                                   'name': 'repache'},
                                                  {'forkCount': 4,
                                                   'issues': {'totalCount': 3},
                                                   'name': 'nf-hishape'},
                                                  {'forkCount': 4,
                                                   'issues': {'totalCount': 7},
                                                   'name': 'ip-countryside'},
                                                  {'forkCount': 42,
                                                   'issues': {'totalCount': 29},
                                                   'name': 'bonesi'},
                                                  {'forkCount': 13,
                                                   'issues': {'totalCount': 3},
                                                   'name': 'rapidminer-anomalydetection'},
                                                  {'forkCount': 0,
                                                   'issues': {'totalCount': 0},
                                                   'name': 'rapidminer-studio'}]}}}}
    
    with open('Markus.csv', 'w') as markus:
        print ('name,forkCount,issues', file=markus)
        for node in Markus['data']['user']['repositories']['nodes']:
            print ('{},{},{}'.format(node['name'], node['forkCount'], node['issues']['totalCount']), file=markus)
    
    • 第一个 print 语句将标题行输出到 csv 文件。
    • for 循环安排从字典中解包项目。
    • 第二个print 语句安排将每个解压缩的项目输出到 csv 文件。

    结果是这样的。

    name,forkCount,issues
    repache,0,0
    nf-hishape,4,3
    ip-countryside,4,7
    bonesi,42,29
    rapidminer-anomalydetection,13,3
    rapidminer-studio,0,0
    

    【讨论】:

    • 也谢谢你。它也有效,但我更喜欢接受的解决方案:)
    【解决方案3】:

    所以考虑到您正在分析 RapidMiner 的使用情况,您也可以选择只使用 RapidMiner 文本处理:

    这是 XML:

    <?xml version="1.0" encoding="UTF-8"?>
    <process version="8.0.001">
      <context>
        <input/>
        <output/>
        <macros/>
      </context>
      <operator activated="true" class="process" compatibility="8.0.001" expanded="true" name="Process">
        <process expanded="true">
          <operator activated="true" class="text:create_document" compatibility="7.5.000" expanded="true" height="68" name="Create Document" width="90" x="45" y="34">
            <parameter key="text" value="{&#10;  &quot;data&quot;: {&#10;    &quot;user&quot;: {&#10;      &quot;name&quot;: &quot;Markus Goldstein&quot;,&#10;      &quot;repositories&quot;: {&#10;        &quot;nodes&quot;: [&#10;          {&#10;            &quot;forkCount&quot;: 0,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 0&#10;            },&#10;            &quot;name&quot;: &quot;repache&quot;&#10;          },&#10;          {&#10;            &quot;forkCount&quot;: 4,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 3&#10;            },&#10;            &quot;name&quot;: &quot;nf-hishape&quot;&#10;          },&#10;          {&#10;            &quot;forkCount&quot;: 4,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 7&#10;            },&#10;            &quot;name&quot;: &quot;ip-countryside&quot;&#10;          },&#10;          {&#10;            &quot;forkCount&quot;: 42,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 29&#10;            },&#10;            &quot;name&quot;: &quot;bonesi&quot;&#10;          },&#10;          {&#10;            &quot;forkCount&quot;: 13,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 3&#10;            },&#10;            &quot;name&quot;: &quot;rapidminer-anomalydetection&quot;&#10;          },&#10;          {&#10;            &quot;forkCount&quot;: 0,&#10;            &quot;issues&quot;: {&#10;              &quot;totalCount&quot;: 0&#10;            },&#10;            &quot;name&quot;: &quot;rapidminer-studio&quot;&#10;          }&#10;        ]&#10;      }&#10;    }&#10;  }&#10;}"
            />
          </operator>
          <operator activated="true" class="text:json_to_data" compatibility="7.5.000" expanded="true" height="82" name="JSON To Data" width="90" x="179" y="34" />
          <connect from_op="Create Document" from_port="output" to_op="JSON To Data" to_port="documents 1" />
          <connect from_op="JSON To Data" from_port="example set" to_port="result 1" />
          <portSpacing port="source_input 1" spacing="0" />
          <portSpacing port="sink_result 1" spacing="0" />
          <portSpacing port="sink_result 2" spacing="0" />
        </process>
      </operator>
    </process>

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-06-03
      • 2016-04-17
      • 1970-01-01
      • 2021-10-02
      • 2016-01-12
      • 1970-01-01
      相关资源
      最近更新 更多