【发布时间】:2021-08-31 04:53:11
【问题描述】:
我有一个字符串类型的 pyspark 列,其中包含一个字典数组。
x = {"a":1,"b":[{"type":"abc","unitValue":"4.4"}]}
我想将字符串转换为 struct 数组,但是在这样做的同时,新列中的字段被填充为 null。
Databricks 运行时 - 8.3(包括 Apache Spark 3.1.1、Scala 2.12)
我的数据框看起来像:
from pyspark.sql.functions import *
from pyspark.sql.types import *
inputSchema = StructType([StructField("a",StringType(),True),
StructField("b",StringType(),True)])
jsonStruct = StructType([StructField("type",StringType(),True),
StructField("unitValue",StringType(),True)])
df = spark.createDataFrame(data =[x],schema = inputSchema).show()
+---+---------------------------+
| a| b |
+---+---------------------------+
| 1|[{type=abc, unitValue=4.4}]|
+---+---------------------------+
df.printSchema()
root
|-- a: string (nullable = true)
|-- b: string (nullable = true)
我正在使用 from_json 函数来实现相同的功能,但值被填充为 null
df1 = df.withColumn("newvalue",from_json(col("b"),jsonStruct,{"mode" : "PERMISSIVE"}))
display(df1)
+---+----------------------------+----------------------------------+
| a| b | newvalue |
+---+----------------------------+----------------------------------+
| 1|[{type=abc, unitValue=xyz}] |{"type": null, "unitValue": null} |
+---+----------------------------+----------------------------------+
谁能帮帮我
【问题讨论】:
标签: pyspark databricks azure-databricks aws-databricks