【问题标题】:spark yarn-cluster mode hdfs io file path configurationspark yarn-cluster模式hdfs io文件路径配置
【发布时间】:2016-11-05 19:35:39
【问题描述】:

我尝试在名称节点服务器上使用 psuedo-dist 模式运行下面的基本 spark wordcount 示例:hadoop 2.6.0

import org.apache.spark.{SparkConf, SparkContext}

object WordCount {
  def main(args: Array[String]){

    //args(0): input file name, args(1): output dir name
    //e.g. hello.txt hello
    val conf = new SparkConf().setAppName("WordCount")
    val sc = new SparkContext(conf)

    val input = sc.textFile(args(0))
    val words = input.flatMap(_.split(" "))
    val counts = words.map((_, 1)).reduceByKey(_ + _)

    counts.saveAsTextFile(args(1))
  }
}

像这样的 start.sh 文件...

$SPARK_HOME/bin/spark-submit \
--master yarn-cluster \
--class com.gmail.hancury.hdfsio.WordCount \
./target/scala-2.10/sparktest_2.10-1.0.jar hello.txt server_hello

当我像

这样写输入文件路径时

hdfs://master:port/path/to/input/hello.txt 或
hdfs:/master:port/path/to/input/hello.txt 或
/path/to/input/hello.txt

自动附加了一些神秘的附加路径

/user/${user.name}/input/

所以,如果我写了像/user/curycu/input/hello.txt 这样的路径,那么应用的路径是这样的:/user/curycu/input/user/curycu/input/hello.txt

因此发生了 fileNotFound 异常。

我想知道那条神奇的道路究竟从何而来……

我检查了name-node服务器的core-site.xml、yarn-site.xml、hdfs-site.xml、mapred-site.xml、spark_env.sh、spark-defaults.conf,但没有任何线索为/user/${user.name}/input

【问题讨论】:

    标签: scala apache-spark configuration hadoop-yarn hadoop2


    【解决方案1】:

    当您不使用程序集 jar (uber jar) 时会发生上述所有错误...

    不是sbt package
    使用sbt assembly

    【讨论】:

      猜你喜欢
      • 2022-01-26
      • 2016-01-18
      • 1970-01-01
      • 2015-08-20
      • 1970-01-01
      • 1970-01-01
      • 2017-03-14
      • 2015-08-02
      • 1970-01-01
      相关资源
      最近更新 更多