【问题标题】:Kafka topic has partitions with leader=-1 (Kafka Leader Election), while node is up and runningKafka 主题具有 leader=-1 的分区(Kafka Leader Election),而节点已启动并正在运行
【发布时间】:2019-07-15 19:32:38
【问题描述】:

我有 3 个成员的 kafka-cluster 设置,__consumer_offsets 主题有 50 个分区。

describe 命令的结果如下:

root@kafka-cluster-0:~# kafka-topics.sh --zookeeper localhost:2181 --describe
Topic:__consumer_offsets    PartitionCount:50   ReplicationFactor:1 Configs:segment.bytes=104857600,cleanup.policy=compact,compression.type=producer
    Topic: __consumer_offsets   Partition: 0    Leader: 1   Replicas: 1 Isr: 1
    Topic: __consumer_offsets   Partition: 1    Leader: -1  Replicas: 2 Isr: 2
    Topic: __consumer_offsets   Partition: 2    Leader: 0   Replicas: 0 Isr: 0
    Topic: __consumer_offsets   Partition: 3    Leader: 1   Replicas: 1 Isr: 1
    Topic: __consumer_offsets   Partition: 4    Leader: -1  Replicas: 2 Isr: 2
    Topic: __consumer_offsets   Partition: 5    Leader: 0   Replicas: 0 Isr: 0
    ...
    ...

成员是节点 0、1 和 2。

很明显,replica=2 中的分区没有为它们设置领导者,并且它们的 leader=-1

想知道这个问题是什么原因造成的,我重启了2nd member kafka服务,没想到会有这个副作用。

而且现在,所有节点都已经运行了几个小时,这是 ls broker/ids 的结果:

/home/kafka/bin/zookeeper-shell.sh localhost:2181 <<< "ls /brokers/ids"
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is disabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[0, 1, 2]

此外,集群中有许多主题,节点 2 不是其中任何一个的领导者,并且它只有数据(replication-factor=1,并且分区托管在此节点),leader=-1,从下面很明显。

Here, node 2 is in the ISR, but never a leader, since replication-factor=2.
Topic:upstream-t2   PartitionCount:20   ReplicationFactor:2 Configs:retention.ms=172800000,retention.bytes=536870912
    Topic: upstream-t2  Partition: 0    Leader: 1   Replicas: 1,2   Isr: 1,2
    Topic: upstream-t2  Partition: 1    Leader: 0   Replicas: 2,0   Isr: 0
    Topic: upstream-t2  Partition: 2    Leader: 0   Replicas: 0,1   Isr: 0
    Topic: upstream-t2  Partition: 3    Leader: 0   Replicas: 1,0   Isr: 0
    Topic: upstream-t2  Partition: 4    Leader: 1   Replicas: 2,1   Isr: 1,2
    Topic: upstream-t2  Partition: 5    Leader: 0   Replicas: 0,2   Isr: 0
    Topic: upstream-t2  Partition: 6    Leader: 1   Replicas: 1,2   Isr: 1,2


Here, node 2 is the only partition some chunks of data are hosted on, but leader=-1.
Topic:upstream-t20  PartitionCount:10   ReplicationFactor:1 Configs:retention.ms=172800000,retention.bytes=536870912
    Topic: upstream-t20 Partition: 0    Leader: 1   Replicas: 1 Isr: 1
    Topic: upstream-t20 Partition: 1    Leader: -1  Replicas: 2 Isr: 2
    Topic: upstream-t20 Partition: 2    Leader: 0   Replicas: 0 Isr: 0
    Topic: upstream-t20 Partition: 3    Leader: 1   Replicas: 1 Isr: 1
    Topic: upstream-t20 Partition: 4    Leader: -1  Replicas: 2 Isr: 2

Any help with how to fix the leader not being elected is greatly appreciated.

另外,很高兴知道这可能对我的经纪人的行为产生任何影响。

编辑---

Kafka 版本:1.1.0 (2.12-1.1.0) 可用空间,例如 800GB 的可用磁盘。 日志文件很正常,在节点 2 上,下面是日志文件的最后 10 行。如果有什么特别要找的,请告诉我。

[2018-12-18 10:31:43,828] INFO [Log partition=upstream-t14-1, dir=/var/lib/kafka] Rolled new log segment at offset 79149636 in 2 ms. (kafka.log.Log)
[2018-12-18 10:32:03,622] INFO Updated PartitionLeaderEpoch. New: {epoch:10, offset:6435}, Current: {epoch:8, offset:6386} for Partition: upstream-t41-8. Cache now contains 7 entries. (kafka.server.epoch.LeaderEpochFileCache)
[2018-12-18 10:32:03,693] INFO Updated PartitionLeaderEpoch. New: {epoch:10, offset:6333}, Current: {epoch:8, offset:6324} for Partition: upstream-t41-3. Cache now contains 7 entries. (kafka.server.epoch.LeaderEpochFileCache)
[2018-12-18 10:38:38,554] INFO [GroupMetadataManager brokerId=2] Removed 0 expired offsets in 0 milliseconds. (kafka.coordinator.group.GroupMetadataManager)
[2018-12-18 10:40:04,831] INFO Updated PartitionLeaderEpoch. New: {epoch:10, offset:6354}, Current: {epoch:8, offset:6340} for Partition: upstream-t41-9. Cache now contains 7 entries. (kafka.server.epoch.LeaderEpochFileCache)
[2018-12-18 10:48:38,554] INFO [GroupMetadataManager brokerId=2] Removed 0 expired offsets in 0 milliseconds. (kafka.coordinator.group.GroupMetadataManager)
[2018-12-18 10:58:38,554] INFO [GroupMetadataManager brokerId=2] Removed 0 expired offsets in 0 milliseconds. (kafka.coordinator.group.GroupMetadataManager)
[2018-12-18 11:05:50,770] INFO [ProducerStateManager partition=upstream-t4-17] Writing producer snapshot at offset 3086815 (kafka.log.ProducerStateManager)
[2018-12-18 11:05:50,772] INFO [Log partition=upstream-t4-17, dir=/var/lib/kafka] Rolled new log segment at offset 3086815 in 2 ms. (kafka.log.Log)
[2018-12-18 11:07:16,634] INFO [ProducerStateManager partition=upstream-t4-11] Writing producer snapshot at offset 3086497 (kafka.log.ProducerStateManager)
[2018-12-18 11:07:16,635] INFO [Log partition=upstream-t4-11, dir=/var/lib/kafka] Rolled new log segment at offset 3086497 in 1 ms. (kafka.log.Log)
[2018-12-18 11:08:15,803] INFO [ProducerStateManager partition=upstream-t4-5] Writing producer snapshot at offset 3086616 (kafka.log.ProducerStateManager)
[2018-12-18 11:08:15,804] INFO [Log partition=upstream-t4-5, dir=/var/lib/kafka] Rolled new log segment at offset 3086616 in 1 ms. (kafka.log.Log)
[2018-12-18 11:08:38,554] INFO [GroupMetadataManager brokerId=2] Removed 0 expired offsets in 0 milliseconds. (kafka.coordinator.group.GroupMetadataManager)

编辑 2 ----

好吧,我已经停止了 leader zookeeper 实例,现在第二个 zookeeper 实例被选为领导者!至此,未选择的领导者问题现已解决!

我不知道可能出了什么问题,所以非常欢迎任何关于“为什么更换 zookeeper 领导者会解决未选择的领导者问题”的想法!

谢谢!

【问题讨论】:

  • 哪个 Kafka 版本?经纪人 2 驱动器上是否有可用空间?日志文件说什么?如果是生产集群,则应考虑增加 __consumer_offsets 的复制因子。
  • 您是否在 server.properties 中明确设置了 broker.id?你是如何安装卡夫卡的?也许经纪人 2 属性文件与其他人可以帮助某人重现问题
  • @SpiXel grep 用于日志文件中的错误。识别控制器并检查它的controller.log,可能是leader选举任务崩溃了。检查其他代理是否仍然可以看到代理 2。
  • @cricket_007 是的,它们在 server.properties 中明确设置。 Kafka 的二进制文件从网站下载并解压到用户的主目录,一切都由提供的脚本运行。配置文件的差异仅限于:broker.idadvertised.listeners。重新启动领导者 zookeeper 服务(更改领导者)实际上修复了选举过程,但我很难弄清楚为什么集群可能会进入该状态。
  • FWIW,我建议增加偏移主题的复制因子

标签: apache-kafka apache-zookeeper kafka-topic


【解决方案1】:

虽然根本原因一直没有确定,但提问者似乎确实找到了解决方案:

我已经停止了领导 zookeeper 实例,现在是第二个 zookeeper 实例被选为领导者!有了这个,未被选中的领导者 问题现已解决!

【讨论】:

    猜你喜欢
    • 2017-05-11
    • 2015-12-25
    • 2013-01-03
    • 2020-11-13
    • 2017-06-20
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-10-01
    相关资源
    最近更新 更多