【问题标题】:Spark - How to change the name of the coalesced parquet fileSpark - 如何更改合并的镶木地板文件的名称
【发布时间】:2018-12-25 11:44:57
【问题描述】:

因此,在将 parquet 文件写入 s3 时,我可以使用以下代码更改目录名称:

spark_NCDS_df.coalesce(1).write.parquet(s3locationC1+"parquet")

现在,当我输出这个时,该目录中的内容如下:

我想做两处改变:

  • 我可以更新part-0000....snappy.parquet 文件的文件名吗?

  • 我可以在没有_SUCCESS、_committed 和_started 文件的情况下输出此文件吗?

我在网上找到的文档并不是很有帮助。

【问题讨论】:

  • Spark SQL 不支持名称自定义(您必须重命名结果),第二个问题看起来像 How to avoid generating crc files and SUCCESS files while saving a DataFrame? 的重复项
  • 命令sc.hadoopConfiguration.set("mapreduce.fileoutputcommitter.marksuccessfuljobs", "false") 给我以下错误:AttributeError: 'RemoteContext' object has no attribute 'hadoopConfiguration'
  • 为什么不只是一个后处理步骤,使用dbutils来清理和重命名?

标签: apache-spark amazon-s3 parquet databricks


【解决方案1】:
    out_file_name = snappy.parquet
    path = "mnt/s3locationC1/"
    tmp_path = "mnt/s3locationC1/tmp_data"
    df = spark_NCDS_df

    def copy_file(path,tmp_path,df,out_file_name):
      df.coalesce(1).write.parquet(tmp_path)
      file = dbutils.fs.ls(tmp_path)[-1][0]
      dbutils.fs.cp(file,path+out_file_name)
      dbutils.fs.rm(tmp_path,True)

   copy_file(path,tmp_path,df,out_file_name)

此函数将您所需的输出文件复制并粘贴到目标位置,然后删除临时文件,所有 _SUCCESS、_committed 和 _started 都随之删除。

如果你还需要什么,请告诉我。

【讨论】:

    猜你喜欢
    • 2021-02-18
    • 1970-01-01
    • 1970-01-01
    • 2019-11-20
    • 1970-01-01
    • 2018-01-21
    • 2019-06-02
    • 1970-01-01
    • 2016-07-04
    相关资源
    最近更新 更多