【发布时间】:2021-10-14 18:07:05
【问题描述】:
我们使用带有 Kafka 集成选项的 Azure 事件中心。我们的服务基于 Java、Spring Boot、Spring Cloud Stream。它们部署在 Azure AKS 上。我们已在集群的虚拟网络上为 Azure 事件中心启用服务终结点。
大多数时候,一切正常。
生产者有时无法发布到 Kafka。我们会丢失消息,这通常对整体数据一致性至关重要。
发生这种情况时,我们会在日志中看到一些错误(为了便于阅读,我将它们分成多行):
日志中的第一个示例:
2019-02-21 22:11:04.681 WARN 1 --- [ad | producer-2]
o.a.k.clients.producer.internals.Sender : [Producer clientId=producer-2]
Got error produce response with correlation id 6 on topic-partition _topic-name_-1,
retrying (4 attempts left). Error: NETWORK_EXCEPTION
第二个例子:
org.apache.kafka.common.errors.TimeoutException:
Expiring 1 record(s) for _topic-name_-1:
30096 ms has passed since batch creation plus linger time
消费者偶尔也会遇到连接问题:
2019-02-22 03:03:59.733 INFO 1 --- [container-0-C-1]
o.a.k.c.c.internals.AbstractCoordinator :
[Consumer clientId=consumer-6, groupId=my-super-service]
Group coordinator my-super-hub.servicebus.windows.net:9093
(id: 2147483647 rack: null) is unavailable or invalid, will attempt rediscovery
是否有人对 Azure 事件中心有类似的问题,或者对可能出现的问题有一些想法?
【问题讨论】:
-
嗨 Nikolaos,你能解决这个问题吗?我目前正在观察类似的问题。
-
嗨 Muton,我已向 Azure 开具了支持票证,并被告知要增加
request.timeout.ms属性的超时时间。这并没有真正帮助,但我正在尝试其他属性,如batch.size...我仍在监视情况,但似乎batch.sizezero 有帮助。仍然不确定它是否真的修复了。 -
它没有帮助。我们已经尝试调整各种设置,但仍然会出现这些错误。我得到了微软支持的回复,当连接空闲很长时间(这是我们的情况)时,这实际上是预期的:“我与我的主题专家团队讨论了这个问题,他们说你遇到的情况是预期的连接在一段时间内处于空闲状态,这是事件中心中的预期行为。”
标签: java azure apache-kafka spring-cloud-stream azure-eventhub