【发布时间】:2018-06-28 12:34:34
【问题描述】:
我正在 HDP 中运行以下工作。
export SPARK-MAJOR-VERSION=2
spark-submit --class com.spark.sparkexamples.Audit --master yarn --deploy-mode cluster \
--files /bigdata/datalake/app/config/metadata.csv BRNSAUDIT_v4.jar dl_raw.ACC /bigdatahdfs/landing/AUDIT/BW/2017/02/27/ACC_hash_total_and_count_20170227.dat TH 20170227
失败并出现以下错误:
*Table or view not found: `dl_raw`.`ACC`; line 1 pos 94;
'Aggregate [count(1) AS rec_cnt#58L, 'count('BRCH_NUM) AS hashcount#59, 'sum('ACC_NUM) AS hashsum#60]
+- 'Filter (('trim('country_code) = trim(TH)) && ('from_unixtime('unix_timestamp('substr('bus_date, 0, 11), MM/dd/yyyy), yyyyMMdd) = 20170227))
+- 'UnresolvedRelation `dl_raw`.`ACC'*
而表存在于 Hive 中,并且可以从 spark-shell 访问。
更新。
val sparkSession = SparkSession.builder
.appName("spark session example")
.enableHiveSupport()
.getOrCreate()
sparkSession.conf.set("spark.sql.crossJoin.enabled", "true")
val df_table_stats = sparkSession.sql("""select count(*) as rec_cnt,count(distinct BRCH_NUM) as hashcount,
sum(ACC_NUM) as hashsum
from dl_raw.ACC
where trim(country_code) = trim('BW')
and from_unixtime(unix_timestamp(substr(bus_date,0,11),'MM/dd/yyyy'),'yyyyMMdd')='20170227'""")
【问题讨论】:
-
分享您的代码?您是否使用 hive 上下文访问它?
-
val sparkSession = SparkSession.builder .appName("spark session example") .enableHiveSupport() .getOrCreate() sparkSession.conf.set("spark.sql.crossJoin.enabled", "true" ) val df_table_stats = sparkSession.sql("select count(*) as rec_cnt,count(distinct BRCH_NUM) as hashcount,sum(ACC_NUM) as hashsum from dl_raw.ACC where trim(country_code) = trim('BW') and from_unixtime( unix_timestamp(substr(bus_date,0,11),'MM/dd/yyyy'),'yyyyMMdd')='20170227'")
-
我使用的是 spark 2,所以我猜 hiveContext 不是强制性的。我得到了以下链接,但不确定如何使用。 community.hortonworks.com/questions/24654/…community.hortonworks.com/questions/5798/…
标签: scala hadoop apache-spark hive apache-spark-sql