【发布时间】:2017-09-22 15:29:16
【问题描述】:
我正在使用 prometheus 来监控 linux 机器上的网络流量。我看到了几个有用的指标,例如 node_network_receive_bytes、node_network_transmit_bytes、drops 和 errs。当我不知道机器的网络带宽时,在 network_in 和 network_out 上设置警报的更好方法是什么?
【问题讨论】:
标签: linux prometheus
我正在使用 prometheus 来监控 linux 机器上的网络流量。我看到了几个有用的指标,例如 node_network_receive_bytes、node_network_transmit_bytes、drops 和 errs。当我不知道机器的网络带宽时,在 network_in 和 network_out 上设置警报的更好方法是什么?
【问题讨论】:
标签: linux prometheus
您应该使用一些查询以获得良好的网络监控结果。 我在 Grafana 上使用了一些查询,并与您分享:
查询出站
sum (irate(node_network_transmit_bytes{hostname=~"$hostname", device!~"lo|bond[0-9]|cbr[0-9]|veth.*"}[1m])) by (hostname) > 0图例格式:{{hostname}} - {{device}} - 出站
查询入站
sum (irate(node_network_receive_bytes{hostname=~"$hostname", device!~"lo|bond[0-9]|cbr[0-9]|veth.*"}[1m])) by (hostname) > 0图例格式:{{hostname}} - {{device}} - 入站
eno(或任何您想要的)设备的网络技术:
图例格式:
{{hostname}} - ({{device}})_in愤怒(node_network_receive_bytes{hostname=~'$hostname',device=~"^en.*"}[5m])*8
图例格式:
{{hostname}} - ({{device}})_out愤怒(node_network_transmit_bytes{hostname=~'$hostname',device=~"^en.*"}[5m])*8
netstas:
图例格式:
{{hostname}} establishednode_netstat_Tcp_CurrEstab{hostname=~'$hostname'}
udp 统计:
愤怒(node_netstat_Udp_InDatagrams{hostname=~"$hostname"}[5m])
愤怒(node_netstat_Udp_InErrors{hostname=~"$hostname"}[5m])
愤怒(node_netstat_Udp_OutDatagrams{hostname=~"$hostname"}[5m])
愤怒(node_netstat_Udp_NoPorts{hostname=~"$hostname"}[5m])
合同
图例格式:
Queue Used ({{hostname}})node_nf_conntrack_entries{hostname=~"$hostname"}/node_nf_conntrack_entries_limit{hostname=~"$hostname"}
请注意主机名。它是 Grafan 上的模板变量。而图例格式是 Grafana 上指标的标签解析。
【讨论】: