【问题标题】:How to get max cpu useage of a pod in kubernetes over a time interval (say 30 days) in promql?如何在 promql 中的时间间隔(比如 30 天)内获得 kubernetes 中 pod 的最大 cpu 使用率?
【发布时间】:2019-11-07 11:13:29
【问题描述】:

我正在尝试估计资源 (cpu) 请求和限制值,我想知道最近一个月使用 prometheus 的 pod 的最大 cpu 使用率。

我检查了这个问题,但无法得到我想要的 Generating range vectors from return values in Prometheus queries

我试过了,但似乎 max_over_time 并没有超过费率

max (  
  max_over_time(
    rate(
      container_cpu_usage_seconds_total[5m]
    )[30d]
  )
) by (pod_name)

无效的参数“查询”:字符 64 处的解析错误:范围规范必须以度量选择器开头,但要跟在 *promql.Call 之后

【问题讨论】:

  • 你的 prometheus 版本是多少?

标签: kubernetes prometheus promql prometheus-operator


【解决方案1】:

您需要将内部表达式(容器 cpu 使用率)捕获为 recording rule:

- record: container_cpu_usage_seconds_total:rate5m
  expr: rate(container_cpu_usage_seconds_total[5m])

然后使用这个新的时间序列来计算 max_over_time:

max (  
  max_over_time(container_cpu_usage_seconds_total:rate5m[30d])
) by (pod_name)

这仅在早于 2.7 的 Prometheus 版本中需要 subqueries can be calculated on the fly,请参阅 this blog post for more details。

请记住,如果您打算使用此复合查询(过去 30 天内收集的最大数据的 max_per_time)进行警报或可视化(而不是一次性查询) ),那么您仍希望使用记录规则来提高查询的性能。它是经典的 CS 计算复杂度权衡(将记录规则存储为单独的时间序列所需的内存/存储空间与处理 30 天数据所需的计算资源!)

【讨论】:

    【解决方案2】:

    请尝试以下方法:

    max_over_time(sum(rate(container_cpu_usage_seconds_total{pod="pod-name-here-759b8f",container_name!="POD", container_name!=""}[1m])) [720h:1s])

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2023-03-12
      • 2020-03-13
      • 2019-07-26
      • 1970-01-01
      • 2018-12-18
      • 2017-02-27
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多