【问题标题】:Save schema of dataframe in S3 location将数据框的架构保存在 S3 位置
【发布时间】:2023-03-04 16:54:02
【问题描述】:

我使用以下命令从 S3 中的数据块中读取了镶木地板文件

df = sqlContext.read.parquet('s3://path/to/parquet/file')

我想读取数据框的架构,可以使用以下命令:

df_schema = df.schema.json()

但我无法将 df_schama 对象写入 S3 上的文件。 注意:我愿意不创建 json 文件。我只想将数据框的架构保存到 AWS S3 中的任何文件类型(可能是文本文件)。

我已经尝试如下编写 json 架构,

df_schema.write.csv("s3://path/to/file")

或

a.write.format('json').save('s3://path/to/file')

他们都给我以下错误:

AttributeError: 'str' object has no attribute 'write'

【问题讨论】:

    标签: apache-spark amazon-s3 pyspark databricks


    【解决方案1】:

    df.schema.json() 结果 string 对象和 string 对象将没有 .write 方法。

    In RDD Api:

    df_schema = df.schema.json()
    

    并行化df_schema变量创建rdd,然后使用.saveAsTextFile方法将schema写入s3。

    sc.parallelize([df_schema]).saveAsTextFile("s3://path/to/file")
    

    (或)

    In Dataframe Api:

    from pyspark.sql import Row
    df_schema = df.schema.json()
    df_sch=sc.parallelize([Row(schema=df1)]).toDF()
    df_sch.write.csv("s3://path/to/file")
    df_sch.write.text("s3://path/to/file") //write as textfile
    

    【讨论】:

      【解决方案2】:

      这是一个保存架构并将其应用于新 csv 数据的工作示例:

      # funcs
      from pyspark.sql.functions import *
      from pyspark.sql.types import *
      
      # example old df schema w/ long datatype
      df = spark.range(10)
      df.printSchema()
      df.write.mode("overwrite").csv("old_schema")
      
      root
       |-- id: long (nullable = false)
      
      # example new df schema we will save w/ int datatype
      df = df.select(col("id").cast("int"))
      df.printSchema()
      
      root
       |-- id: integer (nullable = false)
      
      # get schema as json object
      schema = df.schema.json()
      
      # write/read schema to s3 as .txt
      import json
      
      with open('s3:/path/to/schema.txt', 'w') as F:  
          json.dump(schema, F)
      
      with open('s3:/path/to/schema.txt', 'r') as F:  
          saved_schema = json.load(F)
      
      # saved schema
      saved_schema
      '{"fields":[{"metadata":{},"name":"id","nullable":false,"type":"integer"}],"type":"struct"}'
      
      # construct saved schema object
      new_schema = StructType.fromJson(json.loads(saved_schema))
      
      new_schema
      StructType(List(StructField(id,IntegerType,false)))
      
      # use saved schema to read csv files ... new df has int datatype and not long
      new_df = spark.read.csv("old_schema", schema=new_schema)
      new_df.printSchema()
      root
       |-- id: integer (nullable = true)
      
      

      【讨论】:

      • 虽然代码不错,但是如果 json 类型是 boolean 或 dict 则不起作用。也有存储单引号等错误。
      猜你喜欢
      • 1970-01-01
      • 2021-10-11
      • 1970-01-01
      • 2021-10-16
      • 2012-04-06
      • 2019-07-06
      • 1970-01-01
      • 2018-02-02
      • 1970-01-01
      相关资源
      最近更新 更多