【问题标题】:Debug random SIGSEGV crash调试随机 SIGSEGV 崩溃
【发布时间】:2019-11-03 04:14:05
【问题描述】:

在分布式模式下运行 Kafka Connect 应用程序时,我们遇到了随机的 JVM 崩溃。连接应用程序运行具有自定义任务实现的自定义连接器。该应用程序在以 Alpine Linux 作为基础映像的 Docker 容器上运行。崩溃是完全随机的,我的意思是:

  1. 错误日志没有指向每次崩溃的相同堆栈跟踪(见下文)
  2. 崩溃发生在不同机器上的随机时间点
  3. 推送底层 VM(高 CPU 负载、高内存负载、高 IO 磁盘负载)对崩溃频率没有任何影响

Machine 1 Crash

# A fatal error has been detected by the Java Runtime Environment:
#
#  SIGSEGV (0xb) at pc=0x00007fc1e92df777, pid=48, tid=0x00007fc1d899eae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# V  [libjvm.so+0x4c4777]  JVM_FindSignal+0x52586
#
# Core dump written. Default location: /home/kafka/core or core.48

Machine 2 Crash

# A fatal error has been detected by the Java Runtime Environment:
#
#  SIGSEGV (0xb) at pc=0x00007f7825f1da82, pid=48, tid=0x00007f78179bbae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# j  io.prometheus.jmx.shaded.io.prometheus.client.exporter.common.TextFormat.write004(Ljava/io/Writer;Ljava/util/Enumeration;)V+115
#
# Core dump written. Default location: /home/kafka/core or core.48

Machine 3 Crash

# A fatal error has been detected by the Java Runtime Environment:
#
#  SIGSEGV (0xb) at pc=0x00007fe4463fc2bc, pid=48, tid=0x00007fe435e06ae8
#
# JRE version: OpenJDK Runtime Environment (8.0_151-b12) (build 1.8.0_151-b12)
# Java VM: OpenJDK 64-Bit Server VM (25.151-b12 mixed mode linux-amd64 compressed oops)
# Derivative: IcedTea 3.6.0
# Distribution: Custom build (Tue Nov 21 11:22:36 GMT 2017)
# Problematic frame:
# C  [libjvm.so+0x27b2bc]
#
# Core dump written. Default location: /home/kafka/core or core.48

Machine 4 Crash

# A fatal error has been detected by the Java Runtime Environment:
#
#  SIGSEGV (0xb) at pc=0x00007f130fcb1b93, pid=48, tid=0x00007f130d65a700
#
# JRE version: OpenJDK Runtime Environment (8.0_212-b03) (build 1.8.0_212-8u212-b03-2~deb9u1-b03)
# Java VM: OpenJDK 64-Bit Server VM (25.212-b03 mixed mode linux-amd64 compressed oops)
# Problematic frame:
# V  [libjvm.so+0x760b93]
#
# Core dump written. Default location: /home/kafka/core or core.48

列表继续这样。 其他一些需要提及的事情:

  • 没有应用程序日志
  • 没有写入核心转储(检查了错误文件中提到的位置,但那里没有任何内容)

到目前为止我们尝试过的事情都没有效果:

  • 从基于 Alpine 的 docker 映像切换到 Debian
  • 不包括 Prometheus 代理
  • 将 Open JDK 版本从 8.0.151 更新到 8.0.212

任何有关发现问题的提示将不胜感激!

【问题讨论】:

  • 系统是否启用了核心转储捕获?
  • 看起来像堆损坏。首先,检查问题是否是由JVM本身引起的。尝试 1) 不同的 GC,例如-XX:+UseConcMarkSweepGC; 2)禁用分层编译:-XX:-TieredCompilation; 3) 只保留 C1 编译器:-XX:TieredStopAtLevel=1。顺便说一句,要启用核心转储,请使用 ulimit -c unlimited
  • 核心转储已启用,但限制为 500.000。感谢您的建议,目前我们正在使用 Open JDK 11 进行测试,目前似乎很稳定。我会及时通知您。

标签: java linux docker jvm apache-kafka-connect


【解决方案1】:

似乎使用 JRE 11 运行应用程序已经解决了这个问题。该项目仍然使用 Java 8 构建,但使用 Java 11 运行它已经停止了崩溃。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2014-06-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2014-10-17
    • 2013-08-23
    相关资源
    最近更新 更多