【发布时间】:2018-09-10 19:06:19
【问题描述】:
我正在尝试使用 elasticsearch-hadoop 连接器在 ElasticSearch 中为以下架构的 DataFrame 编制索引。
|-- ROW_ID: long (nullable = false)
|-- SUBJECT_ID: long (nullable = false)
|-- HADM_ID: long (nullable = true)
|-- CHARTDATE: date (nullable = false)
|-- CATEGORY: string (nullable = false)
|-- DESCRIPTION: string (nullable = false)
|-- CGID: integer (nullable = true)
|-- ISERROR: integer (nullable = true)
|-- TEXT: string (nullable = true)
当将此 DataFrame 写入 ElasticSearch 时,“CHARTDATE”字段被写入为 long。根据我正在使用的连接器的文档(如下所示),Spark 中的DateType 字段应在 ElasticSearch 中写为字符串格式的日期。由于我希望利用日期字段在 Kibana 中构建一些可视化,因此将它们写成 long 被证明是有问题的。
https://www.elastic.co/guide/en/elasticsearch/hadoop/6.4/spark.html
用于产生错误的代码
val elasticOptions = Map(
"es.nodes" -> esIP,
"es.port" -> esPort,
"es.mapping.id" -> primaryKey,
"es.index.auto.create" -> "yes",
"es.nodes.wan.only" -> "true",
"es.write.operation" -> "upsert",
"es.net.http.auth.user" -> esUser,
"es.net.http.auth.pass" -> esPassword,
"es.spark.dataframe.write.null" -> "true",
"es.mapping.date.rich" -> "true"
)
castedDF.saveToEs(index, elasticOptions)
我是否缺少将这些值写为 ES 日期的步骤?
【问题讨论】:
标签: apache-spark elasticsearch