【问题标题】:concurrent append to hdfs file in spark并发追加到spark中的hdfs文件
【发布时间】:2017-06-24 23:57:06
【问题描述】:

我得到 ex that failed to append_file file is busy hdfs_non_map_reduce

我通过 spark 从 kafka 获取记录并将其放入 cassandra 和 hdfs stream.map(somefunc).saveToCassandra

stream.map(somefunc).foreachRDD(rdd => 
fs.append.write(rdd.collect.mkstring.getBytes)
fs.close)

hdfs 中的复制因子为 1,我使用一个节点集群 spark 具有 2 个工作人员的独立集群

我不想要rdd.toDF.save("append"),因为它会生成很多文件。 有任何想法吗。 或者可能是 hdfs 有方法检查,如果文件正忙于另一个任务?

【问题讨论】:

    标签: hadoop apache-spark hdfs


    【解决方案1】:

    我不想要 rdd.toDF.save("append") 因为它会生成很多文件

    使用rdd.repartition(1).toDF.save("append")将输出文件的数量减少到1个

    【讨论】:

      【解决方案2】:

      这对我也不好,它为每个 rdd 制作文件,但我想要一个大文件一小时或一天

      所以现在我在集群上使用 try catch finally 方案

      try {
      fs.append.write(rdd.collect.mkstring.getBytes)
      }
      catch {
      case ex: IOException => fs.wait(1000)
      }
      finally {
      fs.close
      }
      

      但我认为我有例外,但它工作正常,我将 100k msg 写入 kafka 和 hdfs 上的文件也有,这样我控制它,但我想,这样,如果 ex,msgs不写了,fs.close

      【讨论】:

      • 我明白了,这样保存到 cassandra 是工作人员,但保存到 hdfs 是驱动程序。为什么?
      猜你喜欢
      • 2017-06-26
      • 2021-01-25
      • 2011-11-25
      • 2021-10-19
      • 1970-01-01
      • 2011-09-17
      • 2023-03-16
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多