【发布时间】:2018-07-10 17:50:04
【问题描述】:
为了在我的项目中使用结构化流,我正在我的 hortonworks 2.6.3 环境中测试 spark 2.2.0 和 Kafka 0.10.1 与 Kerberos 的集成,我正在运行下面的示例代码来检查集成。我可以在 IntelliJ 上以 spark 本地模式运行以下程序,没有任何问题,但是当在 Hadoop 集群上移动到纱线集群/客户端模式时,相同的程序会抛出异常。
我知道我可以为 group-id 配置 kafka acl,但是 spark 结构化流会为每个查询生成新的 group-id,因此我无法在 kafka acl 中配置 group-id 以摆脱授权异常。我很善良现在卡住了。
14:19:59 org.apache.spark.sql.streaming.StreamingQueryException: Not authorized to access group: spark-kafka-source-632450e3-a111-4d09-8704-85320c572aeb--1213729126-driver-2
例外:
18/01/31 14:46:34 INFO AbstractLogin: Successfully logged in.
18/01/31 14:46:34 INFO KerberosLogin: TGT refresh thread started.
18/01/31 14:46:34 INFO KerberosLogin: TGT valid starting at: Wed Jan 31 13:51:11 UTC 2018
18/01/31 14:46:34 INFO KerberosLogin: TGT expires: Wed Jan 31 23:51:14 UTC 2018
18/01/31 14:46:34 INFO KerberosLogin: TGT refresh sleeping until: Wed Jan 31 21:58:11 UTC 2018
Exception in thread "main" 18/01/31 14:46:34 INFO AppInfoParser: Kafka version : 0.10.1.2.6.3.0-235
18/01/31 14:46:34 INFO AppInfoParser: Kafka commitId : ba0af6800a08d2f8
org.apache.spark.sql.streaming.StreamingQueryException: Not authorized to access group: spark-kafka-source-632450e3-a111-4d09-8704-85320c572aeb--1213729126-driver-2
=== Streaming Query ===
Identifier: [id = 64a8dbd2-c674-43f7-947d-9aac1667b2b0, runId = 70ce5ee9-ead6-44eb-a7cd-93619b10b811]
Current Committed Offsets: {}
Current Available Offsets: {}
Current State: ACTIVE
Thread State: RUNNABLE
Logical Plan:
Project [value#16]
+- Project [cast(key#0 as string) AS key#15, cast(value#1 as string) AS value#16]
+- StreamingExecutionRelation KafkaSource[Subscribe[test_topic]], [key#0, value#1, topic#2, partition#3, offset#4L, timestamp#5, timestampType#6]
at org.apache.spark.sql.execution.streaming.StreamExecution.org$apache$spark$sql$execution$streaming$StreamExecution$$runBatches(StreamExecution.scala:343)
at org.apache.spark.sql.execution.streaming.StreamExecution$$anon$1.run(StreamExecution.scala:206)
Caused by: org.apache.kafka.common.errors.GroupAuthorizationException: Not authorized to access group: spark-kafka-source-632450e3-a111-4d09-8704-85320c572aeb--1213729126-driver-2
18/01/31 14:46:34 ERROR StreamExecution: Query [id = 01bd97ea-6d2c-446c-a366-491d252925aa, runId = cc8dc932-9297-47c5-b30b-007624c03163] terminated with error
org.apache.kafka.common.errors.GroupAuthorizationException: Not authorized to access group: spark-kafka-source-d690d270-7092-4aed-82c2-97fdfd80d0ed--604732661-driver-2
18/01/31 14:46:34 WARN KerberosLogin: TGT renewal thread has been interrupted and will exit.
18/01/31 14:46:34 INFO SparkContext: Invoking stop() from shutdown hook
18/01/31 14:46:34 INFO AbstractConnector: Stopped Spark@37524c9b{HTTP/1.1,[http/1.1]}{0.0.0.0:4040}
18/01/31 14:46:34 INFO SparkUI: Stopped Spark web UI at http://192.168.0.19:4040
18/01/31 14:46:34 INFO YarnClientSchedulerBackend: Interrupting monitor thread
18/01/31 14:46:34 INFO YarnClientSchedulerBackend: Shutting down all executors
18/01/31 14:46:34 INFO YarnSchedulerBackend$YarnDriverEndpoint: Asking each executor to shut down
【问题讨论】:
-
这个问题有解决办法吗?我有同样的问题。
-
@galmeriol 我还没有尝试过,但似乎这个问题已在 spark 2.3.0 中修复。请在此处检查并更新您的发现。
-
我已经在 Spark 2.3.0 中尝试过,仍然没有运气。我想知道在代理端使用一些通配符解决方案是否合理或可能,例如使用 spark-kafka-* 动态授予组 ID 权限。
标签: hadoop apache-spark apache-kafka kerberos kafka-consumer-api