【问题标题】:How to flatten this json as a tsv?如何将此 json 展平为 tsv?
【发布时间】:2020-01-18 13:01:03
【问题描述】:

我想将此 JSON 扁平化为 tsv 文件。

https://www.vi4io.org/assets/io500/2019-06/data.json

问题在于每个条目({} 的第一级)都有许多字段/子字段。我不想指定这么多的字段名称。并且不能保证所有条目中的字段/子字段都相同。因此,我希望结果列包含所有文件/子字段的联合。列名的排序应尽可能接近原始 json 文件。 (例如,同一字段中的那些子字段应在 tsv 中一起列出)。

将此 json 文件转换为 tsv 的最佳方法是什么?谢谢。

【问题讨论】:

  • 包含您输入的摘录,而不是阴暗的链接。并且输出示例也将非常有帮助
  • user1424739 - 我没有否决你的问题,但如果你想避免进一步的否决,首先按照@OguzIsmail 的建议做(提供一个非常简短的例子来说明关键点您),然后至少展示您使用您可以接受的工具进行的一次尝试。

标签: json csv export-to-csv jq flatten


【解决方案1】:

这是一个 jq 解决方案,它适用于任何 JSON 对象数组,没有限制,但请参阅下面的“注意事项”。

json2tsv.jq

# Given an array of JSON objects, 
# produce "TSV" rows, with a header row.
# Handle terminal arrays specially if they are flat.


# emit a stream
def json2headers:
  def isscalar: type | . != "array" and . != "object";
  def isflat: all(.[]; isscalar);
  paths as $p
  | getpath($p)
  | if type == "array" and isflat then $p
     elif isscalar and (($p[-1]|type) == "string") then $p
     else empty end ;

def json2array($header):
   [$header[] as $p | (try getpath($p) catch null)] ;

def json2tsv:
  ( [.[] | json2headers] | unique) as $h
  | ([$h[]|join("_") ],
     (.[]
      | json2array($h)
      | map( if type == "array" then map(tostring)|join("|") else tostring end)))
  | @tsv ;

用法

jq -r -L. 'include "json2tsv"; json2tsv' input.json

输出

输入的样本非常大,这里我只显示标题,并附上一个单独的例子。

标题

find_easy   information_URL information_client_kernel_version   information_client_nodes    information_client_operating_system information_client_operating_system_version information_client_procs_per_node   information_comment information_data    information_ds_network  information_ds_nodes    information_ds_operating_system_version information_ds_software_version information_ds_storage_devices  information_ds_storage_interface    information_ds_storage_type information_ds_volatile_memory_capacity information_embargo_end_date    information_filesystem_name information_filesystem_type information_filesystem_version  information_id  information_institution information_list    information_md_network  information_md_nodes    information_md_operating_system_version information_md_software_version information_md_storage_devices  information_md_storage_interface    information_md_storage_type information_md_volatile_memory_capacity information_note    information_storage_install_date    information_storage_refresh_date    information_storage_vendor  information_submission_date information_submitter   information_system  information_vendorURL   information_whatever    io500_md    io500_score ior_easy_read   ior_easy_write  ior_hard_read   ior_hard_write  mdtest_easy_delete  mdtest_easy_stat    mdtest_easy_write   mdtest_hard_delete  mdtest_hard_read    mdtest_hard_stat    mdtest_hard_write

简短示例

input.json

[ {a: [1,2], b: {c:3, d: [{e:4},{e:5, f:6}]}},
  {b: {d: [{e:4},{f:6, e:5}], c:3}, a:[101,102] } ]
输出
a   b_c b_d_0_e b_d_1_e b_d_1_f
1|2 3   4   5   6
101|102 3   4   5   6

input.json(Dmitry 的变体)

[ {a:[1,2],b:{c:3,d:[{e:4},{e:5,f:6}]}},
  {b:{d:[{e:4},{f:6}],c:3},a:[101,102]} ]
输出
a   b_c b_d_0_e b_d_1_e b_d_1_f
1|2 3   4   5   6
101|102 3   4   null    6

具有不同结构的对象

[ {a: [1,2], b: {c: 3}},
  {a: [4,5], b: {c: {d: 6 } } }
输出
a   b_c b_c_d
1|2 3   null
4|5 {"d":6} 6

注意事项

  • 对于顶级数组中的每个对象,计算所有标量和标量值数组的路径;如果任何此类路径在另一个顶级对象中无效,则输出中的相应值将为null,如上一个示例所示。

  • 平面数组被转换为管道分隔的值,这样如果输入中包含["1|2", ["3|4"]之类的数组,它将与字符串值无法区分, “1|2|3|4”等。如果这是个问题,当然可以更改用作数组项分隔值的字符。

  • 标题名称可能会发生类似的冲突。

  • jq 的@tsv 为"" 和null 生成一个空字符串,因此如果区分两者很重要,您可能希望在调用@tsv 之前考虑使用适当的map .

【讨论】:

  • 虽然 jq 不关我的事,我已经赞成你的解决方案,知道你有多勤奋,请你更新它,以便它也能与 不规则的 jsons 喜欢:[{a:[1,2],b:{c:3,d:[{e:4},{e:5,f:6}]}},{b:{d:[{e:4},{f:6}],c:3},a:[101,102]}] - 让它更全面? - 会是一次很好的学习经历
  • @Dmitry - 好收获!更新。谢谢。顺便说一句,您的standard.json 是否已经在某处可用?如果没有,请您提供它吗?
  • 让我想想最好的托管方式,一旦完成,我会提供一个链接。此外,我在我的新 Mac 上看到 jq 和 jtc 的性能数字都有所提高,因此,我也会发布其规范以供参考。完成后会更新您。
  • @Dmitry - 再次修改以删除对输入对象的所有限制,但代价是可能在 TSV 输出中包含整个对象(如新的(最后一个)示例所示)。
  • 我已发布 standard.json 并更新了性能数据(并发布了 comp.spec)
【解决方案2】:

如果 jq 不是强制性的,您可以使用 Miller https://github.com/johnkerl/miller

命令是

curl "https://www.vi4io.org/assets/io500/2019-06/data.json" | mlr --j2t unsparsify >output.tsv

如果您想将输出字段名称从 information:data 重命名为 information_data,您可以通过这种方式使用 rename 动词:

curl "https://www.vi4io.org/assets/io500/2019-06/data.json" | mlr --j2t unsparsify then rename -r '(.+):(.+),\1_\2' >output.tsv

【讨论】:

  • 如何将标题中的information:data改为information_data?
  • @user1424739 我已在回复中插入相关的最后说明
猜你喜欢
  • 2021-08-09
  • 1970-01-01
  • 2021-07-16
  • 1970-01-01
  • 2021-04-19
  • 1970-01-01
  • 1970-01-01
  • 2012-07-05
  • 1970-01-01
相关资源
最近更新 更多