【发布时间】:2023-03-04 16:54:02
【问题描述】:
我使用以下命令从 S3 中的数据块中读取了镶木地板文件
df = sqlContext.read.parquet('s3://path/to/parquet/file')
我想读取数据框的架构,可以使用以下命令:
df_schema = df.schema.json()
但我无法将 df_schama 对象写入 S3 上的文件。
注意:我愿意不创建 json 文件。我只想将数据框的架构保存到 AWS S3 中的任何文件类型(可能是文本文件)。
我已经尝试如下编写 json 架构,
df_schema.write.csv("s3://path/to/file")
或
a.write.format('json').save('s3://path/to/file')
他们都给我以下错误:
AttributeError: 'str' object has no attribute 'write'
【问题讨论】:
标签: apache-spark amazon-s3 pyspark databricks