【问题标题】:how to flatten complex nested json in spark dataframe using java dynamically如何动态使用java在spark数据框中展平复杂的嵌套json
【发布时间】:2020-06-06 03:56:44
【问题描述】:

我的输入 json 数据框文件如下所示:

company: array (nullable = true)
|    |-- element: struct (containsNull = true)
|    |    |-- address: struct (nullable = true)
|    |    |    |-- city: string (nullable = true)
|    |    |    |-- county: string (nullable = true)
|    |    |    |-- latitude: string (nullable = true)
|    |    |    |-- line1: string (nullable = true)
|    |    |    |-- line2: string (nullable = true)
|    |    |    |-- longitude: string (nullable = true)
|    |    |    |-- postalCode: long (nullable = true)
|    |    |    |-- state: struct (nullable = true)
|    |    |    |    |-- code: int (nullable = true)
|    |    |    |    |-- name: string (nullable = true)
|    |    |    |-- stateOtherDescription: string (nullable = true)
|    |    |-- addressSourceOther: string (nullable = true)
|    |    |-- addressSourceType: struct (nullable = true)
|    |    |     |-- code: int (nullable = true)
|    |    |     |-- name: string (nullable = true)
|    |    |-- reasons: array (nullable = true)
|    |    |     |-- element: struct (containsNull = true)
|    |    |     |-- improve: string (nullable = true)
|    |    |     |-- far: string (nullable = true)
|    |    |     |-- home: string (nullable = true)

我想使用 spark java 动态地展平它。有人可以帮我解决这个问题吗

【问题讨论】:

  • 我认为您的问题与 apache-spark 无关。输出应该是什么样子?
  • 我在 HDFS 中接收数据并在 spark df(2.3 版)中读取数据。输出应该是单个列中的所有属性,没有嵌套。

标签: java json apache-spark


【解决方案1】:

查看http://stefanfrings.de/bfUtilities/bfUtilities.zip 中的 de.stefanfrings/parsing/JsonSlurper 类。它读取 JSON 文档并根据内容创建平面 HashMap。

【讨论】:

    猜你喜欢
    • 2021-07-14
    • 2023-03-03
    • 1970-01-01
    • 1970-01-01
    • 2023-03-07
    • 2019-12-18
    • 2019-03-18
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多