【问题标题】:Horizontal pod Autoscaler scales custom metric too aggressively on GKE水平 pod Autoscaler 在 GKE 上过于激进地扩展自定义指标
【发布时间】:2020-01-12 03:51:27
【问题描述】:

我在 Google Kubernetes Engine 上有以下 Horizo​​ntal Pod Autoscaler 配置,可通过自定义指标扩展部署 - RabbitMQ messages ready count 用于特定队列:foo-queue。

它正确地获取度量值。

当插入 2 条消息时,它会将部署扩展到最多 10 个副本。 我希望它可以扩展到 2 个副本,因为 targetValue 是 1 并且有 2 条消息准备好了。

为什么它会如此积极地扩展?

HPA 配置:

apiVersion: autoscaling/v2beta1
kind: HorizontalPodAutoscaler
metadata:
  name: foo-hpa
  namespace: development
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: foo
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metricName: "custom.googleapis.com|rabbitmq_queue_messages_ready"
      metricSelector:
        matchLabels:
          metric.labels.queue: foo-queue
      targetValue: 1

【问题讨论】:

  • 你确定targetValue: 1?为什么这个值这么小?我看到推荐值超过 100 的样本
  • @Yasen 当设置 targetValue: 100 并在队列中有 2 条消息时,HPA 扩展到 2 个 pod,这似乎非常激进,我希望它是 1 个副本
  • 请您阅读前 Docker 开发人员 Jérôme Petazzoni 的本指南:Kubernetes Deployments: The Ultimate Guide - Semaphore。它解释了为什么在k8s 中有两个副本,而在docker 中没有一个副本

标签: kubernetes rabbitmq google-kubernetes-engine kubernetes-hpa


【解决方案1】:

我认为 explaining how targetValue works 使用 Horizo​​ntalPodAutoscaler 做得很好。但是,根据您的问题,我认为您正在寻找targetAverageValue 而不是targetValue。

在 the Kubernetes docs on HPAs 中,它提到使用 targetAverageValue 指示 Kubernetes 根据自动缩放器下所有 Pod 公开的平均指标来缩放 Pod。虽然文档没有明确说明,但外部指标(如消息队列中等待的作业数量)计为单个数据点。通过使用 targetAverageValue 对外部指标进行缩放,您可以创建一个自动缩放器来缩放 Pod 的数量以匹配 Pod 与作业的比率。

回到你的例子:

apiVersion: autoscaling/v2beta1
kind: HorizontalPodAutoscaler
metadata:
  name: foo-hpa
  namespace: development
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: foo
  minReplicas: 1
  maxReplicas: 10
  metrics:
  - type: External
    external:
      metricName: "custom.googleapis.com|rabbitmq_queue_messages_ready"
      metricSelector:
        matchLabels:
          metric.labels.queue: foo-queue
      # Aim for one Pod per message in the queue
      targetAverageValue: 1

将导致 HPA 尝试为队列中的每条消息保留一个 Pod(最多 10 个 Pod)。

顺便说一句,每条消息定位一个 Pod 可能会导致您不断地启动和停止 Pod。如果您最终启动了大量 Pod 并处理队列中的所有消息,Kubernetes 会将您的 Pod 缩减为 1。根据启动 Pod 所需的时间以及处理您的消息所需的时间,您可能通过指定更高的targetAverageValue 来降低平均消息延迟。理想情况下,在流量恒定的情况下,您的目标应该是让 Pod 处理消息的数量恒定(这要求您以与消息入队大致相同的速率处理消息)。

【讨论】:

    【解决方案2】:

    根据https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/

    从最基本的角度来看,Horizo​​ntal Pod Autoscaler 控制器根据所需指标值与当前指标值之间的比率进行操作:

    desiredReplicas = ceil[currentReplicas * ( currentMetricValue / desiredMetricValue )]
    

    从上面我了解到,只要队列有消息,k8 HPA 就会继续扩大,因为currentReplicas 是desiredReplicas 计算的一部分。

    例如,如果:

    currentReplicas = 1

    currentMetricValue / desiredMetricValue = 2/1

    然后:

    desiredReplicas = 2

    如果指标在下一个 hpa 周期中保持不变,currentReplicas 将变为 2,desiredReplicas 将提高到 4

    【讨论】:

    • 就是这样。 HPA 不断尝试扩大规模以将您的指标降低到目标值,但它不能,因为 pod 数与队列中的消息数之间没有比率。这是 HPA 自定义指标的常见缺陷
    【解决方案3】:

    尝试按照此说明在k8s 中描述RabbitMQ 的水平自动缩放设置

    Kubernetes Workers Autoscaling based on RabbitMQ queue size

    特别推荐 targetValue: 20 的指标 rabbitmq_queue_messages_ready 而不是 targetValue: 1:

    apiVersion: autoscaling/v2beta1
    kind: HorizontalPodAutoscaler
    metadata:
      name: workers-hpa
    spec:
      scaleTargetRef:
        apiVersion: apps/v1beta1
        kind: Deployment
        name: my-workers
      minReplicas: 1
      maxReplicas: 10
      metrics:
      - type: External
        external:
          metricName: "custom.googleapis.com|rabbitmq_queue_messages_ready"
          metricSelector:
            matchLabels:
              metric.labels.queue: myqueue
          **targetValue: 20
    

    如果 RabbitMQ 队列 myqueue 总共有超过 20 个未处理的作业,我们的部署 my-workers 将会增加

    【讨论】:

    • 问题是指标不会根据.pods的数量而改变。对于 1 个 pos,队列中有 1 条消息,对于 20 个 pod,队列中仍然有 1 条消息。 HPA 正在尝试扩大 pod 的数量以减少当前指标。
    【解决方案4】:

    我正在使用来自 RabbitMQ 的相同 Prometheus 指标(我使用 Celery 和 RabbitMQ 作为代理)。

    这里有没有人考虑过使用rabbitmq_queue_messages_unacked 度量而不是rabbitmq_queue_messages_ready?

    问题是,一旦工人拉出消息,rabbitmq_queue_messages_ready 就会减少,我担心长时间运行的任务可能会被 HPA 杀死,而 rabbitmq_queue_messages_unacked 会一直保持到任务完成。

    例如,我有一条消息将触发一个新的 pod (celery-worker) 来运行一个需要 30 分钟的任务。 rabbitmq_queue_messages_ready 将随着 pod 的运行而减少,HPA 冷却/延迟将终止 pod。

    编辑:似乎第三个 rabbitmq_queue_messages 是正确的 - 这是未确认和就绪的总和:

    就绪和未确认消息的总和 - 总队列深度

    documentation

    【讨论】:

    • 有趣的一点,它提出了一个问题,对于运行 30 分钟的 pod,应该采用什么自动缩放策略。我想这真的取决于业务需求。
    猜你喜欢
    • 2018-03-01
    • 2023-01-13
    • 1970-01-01
    • 2021-11-26
    • 1970-01-01
    • 1970-01-01
    • 2016-12-18
    • 2021-10-31
    • 2022-06-14
    相关资源
    最近更新 更多