【问题标题】:Elastic search master disaster recoveryElastic Search Master 容灾
【发布时间】:2017-01-29 21:17:30
【问题描述】:

我们有一个包含 5 个数据节点和 2 个主节点的弹性搜索集群。一个主节点上的弹性搜索服务始终处于禁用状态,因此始终只有一个主节点处于活动状态。今天由于某种原因,当前的主节点宕机了。我们在第二个主节点上启动了服务。所有连接到新主节点的数据节点,所有主分片均已成功分配,但未分配所有副本,我剩下将近 384 个未分配分片。

我现在应该怎么做,分配他们?

在这种情况下必须执行的最佳做法和步骤是什么?

以下是我的http://es-master-node:9200/_settings 的样子:http://pastebin.com/mK1QBfP6

当我尝试手动分配分片时,我收到以下错误:

➜  Desktop curl -XPOST http://localhost:9200/_cluster/reroute\?pretty -d '{
  "commands": [
    {
      "allocate": {
        "index": "logstash-1970.01.18",
        "shard": 1,
        "node": "node-name",
        "allow_primary": true
      }
    }
  ]
}'
{
  "error" : {
    "root_cause" : [ {
      "type" : "illegal_argument_exception",
      "reason" : "[allocate] allocation of [logstash-1970.01.18][1] on node {node-name}{vrVG4CBbSvubWHOzn2qfQA}{10.100.0.146}{10.100.0.146:9300}{master=false} is not allowed, reason: [YES(allocation disabling is ignored)][NO(more than allowed [85.0%] used disk on node, free: [13.671127301258165%])][YES(shard not primary or relocation disabled)][YES(target node version [2.2.0] is same or newer than source node version [2.2.0])][YES(no allocation awareness enabled)][YES(shard is not allocated to same node or host)][YES(allocation disabling is ignored)][YES(below shard recovery limit of [2])][YES(total shard limit disabled: [index: -1, cluster: -1] <= 0)][YES(node passes include/exclude/require filters)][YES(primary is already active)]"
    } ],
    "type" : "illegal_argument_exception",
    "reason" : "[allocate] allocation of [logstash-1970.01.18][1] on node {node-name}{vrVG4CBbSvubWHOzn2qfQA}{10.100.0.146}{10.100.0.146:9300}{master=false} is not allowed, reason: [YES(allocation disabling is ignored)][NO(more than allowed [85.0%] used disk on node, free: [13.671127301258165%])][YES(shard not primary or relocation disabled)][YES(target node version [2.2.0] is same or newer than source node version [2.2.0])][YES(no allocation awareness enabled)][YES(shard is not allocated to same node or host)][YES(allocation disabling is ignored)][YES(below shard recovery limit of [2])][YES(total shard limit disabled: [index: -1, cluster: -1] <= 0)][YES(node passes include/exclude/require filters)][YES(primary is already active)]"
  },
  "status" : 400
}

任何帮助将不胜感激。

【问题讨论】:

    标签: elasticsearch sharding


    【解决方案1】:

    所以,这是我为分配未分配的分片所做的事情:

    生成 5 个新的 ES-DATA 服务器并等待它们加入集群。一旦他们进入集群,我就使用了以下脚本:

    #!/bin/bash
    array=(node1 node2 node3 node4 node5)
    node_counter=0
    length=${#array[@]}
    IFS=$'\n'
    for line in $(curl -s 'http://ip-adress:9200/_cat/shards'|  fgrep UNASSIGNED); do
        INDEX=$(echo $line | (awk '{print $1}'))
        SHARD=$(echo $line | (awk '{print $2}'))
        NODE=${array[$node_counter]}
        echo $NODE
        curl -XPOST 'http://IP-adress:9200/_cluster/reroute' -d '{
            "commands": [
            {
                "allocate": {
                    "index": "'$INDEX'",
                    "shard": '$SHARD',
                    "node": "'$NODE'",
                    "allow_primary": true
                }
            }
            ]
        }'
        node_counter=$(((node_counter)%length +1))
    done
    

    将未分配的分片分配给新的数据节点。集群再次恢复大约需要 5 到 6 分钟。虽然这是 hack,但相关的答案会更有意义。

    以下是未回答的问题:

    • 旧节点上已经有分片了,为什么 ES-Master 没有意识到这一点?
    • 我们如何明确要求 ES-MASTER 扫描已经存在的数据节点并从中获取信息(关于它们的当前状态、它们拥有的副本、它们包含的分片等)

    【讨论】:

      猜你喜欢
      • 2016-06-01
      • 2019-11-24
      • 2018-01-14
      • 2021-07-31
      • 1970-01-01
      • 2023-03-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多