【发布时间】:2019-08-21 14:21:44
【问题描述】:
我们突然观察到一个表(设备)指标的高写入延迟。
这是一个包含
这是在 RF=3 的 3 节点集群上。 每个节点都有 8GB 内存。我们在 docker 中运行 Cassandra 3.11.4。
日志中没有什么异常。应用程序也运行顺利。
nodetool 表格直方图
Percentile SSTables Write Latency Read Latency Partition Size Cell Count
(micros) (micros) (bytes)
50% 0.00 263.21 0.00 258 17
75% 0.00 1131.75 0.00 372 20
95% 0.00 12108.97 0.00 642 29
98% 0.00 25109.16 0.00 642 35
99% 0.00 43388.63 0.00 642 35
Min 0.00 8.24 0.00 51 0
Max 0.00 155469.30 0.00 770 35
节点工具状态
Datacenter: datacenter-prod
===========================
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
-- Address Load Tokens Owns (effective) Host ID Rack
UN 10.164.0.23 2.62 GiB 256 100.0% e7e2a38a-d4f3-4758-a345-73fcffe26035 rack1
UN 10.164.0.24 2.61 GiB 256 100.0% 0c18b8e4-5ca2-4fb5-9e8c-663b74909fbb rack1
UN 10.164.0.58 2.62 GiB 256 100.0% 547c0746-72a8-4fec-812a-8b926d2426ae rack1
发生了什么事?统计数据是在撒谎还是出现了问题?
编辑: 我能够将问题缩小到其中一个节点。 节点 2 上的导出器显示:
cassandra_stats{cluster="Prod Cluster 2",datacenter="datacenter-prod",keyspace="iot_data",table="devices",name="org:apache:cassandra:metrics:table:iot_data:devices:writelatency:99thpercentile",} 268650.95
而node1和node3是这样的:
cassandra_stats{cluster="Prod Cluster 2",datacenter="datacenter-prod",keyspace="iot_data",table="devices",name="org:apache:cassandra:metrics:table:iot_data:devices:writelatency:99thpercentile",} 10090.808
但我仍然不知道在 node2 上是什么原因造成的。它没有负载,内存使用也很好?!有什么想法吗?
【问题讨论】:
标签: cassandra