【发布时间】:2017-12-22 03:09:45
【问题描述】:
作为 Scala / Spark 的菜鸟,我有点卡住了,希望能得到任何帮助!
正在将 JSON 数据导入 Spark 数据框。在这个过程中,我最终得到了一个与 JSON 输入中存在相同嵌套结构的数据框。
我的目标是使用 Scala 递归地展平整个数据框(包括数组/字典中最内层的子属性)。
此外,可能存在具有相同名称的子属性。因此,也需要区分它们。
此处显示了一个有点相似的解决方案(不同父母的相同子属性) - https://stackoverflow.com/a/38460312/3228300
我希望实现的示例如下:
{
"id": "0001",
"type": "donut",
"name": "Cake",
"ppu": 0.55,
"batters":
{
"batter":
[
{ "id": "1001", "type": "Regular" },
{ "id": "1002", "type": "Chocolate" },
{ "id": "1003", "type": "Blueberry" },
{ "id": "1004", "type": "Devil's Food" }
]
},
"topping":
[
{ "id": "5001", "type": "None" },
{ "id": "5002", "type": "Glazed" },
{ "id": "5005", "type": "Sugar" },
{ "id": "5007", "type": "Powdered Sugar" },
{ "id": "5006", "type": "Chocolate with Sprinkles" },
{ "id": "5003", "type": "Chocolate" },
{ "id": "5004", "type": "Maple" }
]
}
相应的扁平化输出 Spark DF 结构将是:
{
"id": "0001",
"type": "donut",
"name": "Cake",
"ppu": 0.55,
"batters_batter_id_0": "1001",
"batters_batter_type_0": "Regular",
"batters_batter_id_1": "1002",
"batters_batter_type_1": "Chocolate",
"batters_batter_id_2": "1003",
"batters_batter_type_2": "Blueberry",
"batters_batter_id_3": "1004",
"batters_batter_type_3": "Devil's Food",
"topping_id_0": "5001",
"topping_type_0": "None",
"topping_id_1": "5002",
"topping_type_1": "Glazed",
"topping_id_2": "5005",
"topping_type_2": "Sugar",
"topping_id_3": "5007",
"topping_type_3": "Powdered Sugar",
"topping_id_4": "5006",
"topping_type_4": "Chocolate with Sprinkles",
"topping_id_5": "5003",
"topping_type_5": "Chocolate",
"topping_id_6": "5004",
"topping_type_6": "Maple"
}
之前没有过多使用 Scala 和 Spark,不确定如何继续。
最后,如果有人可以请帮助提供通用/非模式解决方案的代码,我将非常感激,因为我需要将它应用于许多不同的集合。
非常感谢:)
【问题讨论】:
标签: json scala apache-spark dataframe flatten