【发布时间】:2019-11-03 04:14:05
【问题描述】:
在分布式模式下运行 Kafka Connect 应用程序时,我们遇到了随机的 JVM 崩溃。连接应用程序运行具有自定义任务实现的自定义连接器。该应用程序在以 Alpine Linux 作为基础映像的 Docker 容器上运行。崩溃是完全随机的,我的意思是:
- 错误日志没有指向每次崩溃的相同堆栈跟踪(见下文)
- 崩溃发生在不同机器上的随机时间点
- 推送底层 VM(高 CPU 负载、高内存负载、高 IO 磁盘负载)对崩溃频率没有任何影响
# A fatal error has been detected by the Java Runtime Environment:
#
# SIGSEGV (0xb) at pc=0x00007fc1e92df777, pid=48, tid=0x00007fc1d899eae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# V [libjvm.so+0x4c4777] JVM_FindSignal+0x52586
#
# Core dump written. Default location: /home/kafka/core or core.48
# A fatal error has been detected by the Java Runtime Environment:
#
# SIGSEGV (0xb) at pc=0x00007f7825f1da82, pid=48, tid=0x00007f78179bbae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# j io.prometheus.jmx.shaded.io.prometheus.client.exporter.common.TextFormat.write004(Ljava/io/Writer;Ljava/util/Enumeration;)V+115
#
# Core dump written. Default location: /home/kafka/core or core.48
# A fatal error has been detected by the Java Runtime Environment:
#
# SIGSEGV (0xb) at pc=0x00007fe4463fc2bc, pid=48, tid=0x00007fe435e06ae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# C [libjvm.so+0x27b2bc]
#
# Core dump written. Default location: /home/kafka/core or core.48
# A fatal error has been detected by the Java Runtime Environment:
#
# SIGSEGV (0xb) at pc=0x00007f130fcb1b93, pid=48, tid=0x00007f130d65a700
#
# JRE version: OpenJDK Runtime Environment (8.0_212-b03) (build 1.8.0_212-8u212-b03-2~deb9u1-b03)
# Java VM: OpenJDK 64-Bit Server VM (25.212-b03 mixed mode linux-amd64 compressed oops)
# Problematic frame:
# V [libjvm.so+0x760b93]
#
# Core dump written. Default location: /home/kafka/core or core.48
列表继续这样。 其他一些需要提及的事情:
- 没有应用程序日志
- 没有写入核心转储(检查了错误文件中提到的位置,但那里没有任何内容)
到目前为止我们尝试过的事情都没有效果:
- 从基于 Alpine 的 docker 映像切换到 Debian
- 不包括 Prometheus 代理
- 将 Open JDK 版本从 8.0.151 更新到 8.0.212
任何有关发现问题的提示将不胜感激!
【问题讨论】:
-
系统是否启用了核心转储捕获?
-
看起来像堆损坏。首先,检查问题是否是由JVM本身引起的。尝试 1) 不同的 GC,例如
-XX:+UseConcMarkSweepGC; 2)禁用分层编译:-XX:-TieredCompilation; 3) 只保留 C1 编译器:-XX:TieredStopAtLevel=1。顺便说一句,要启用核心转储,请使用ulimit -c unlimited -
核心转储已启用,但限制为 500.000。感谢您的建议,目前我们正在使用 Open JDK 11 进行测试,目前似乎很稳定。我会及时通知您。
标签: java linux docker jvm apache-kafka-connect