【问题标题】:Elasticsearch garbage collection warnings (JvmGcMonitorService)Elasticsearch 垃圾收集警告 (JvmGcMonitorService)
【发布时间】:2021-12-11 18:13:14
【问题描述】:

我正在通过 Docker 运行一组集成测试。

测试创建各种索引、索引文档、查询文档、删除文档等。

大多数情况下它们都通过了,但偶尔会有一些测试会失败,因为 SocketTimeout 连接到 Elasticsearch。

发生故障时,我会在日志中看到这些 GC 警告:

2021-08-16T19:27:04.032753121Z 19:27:04.028 elasticsearch> [19:27:04.025Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][young][359][4] duration [7.5s], collections [1]/[8.3s], total [7.5s]/[7.7s], memory [310.9mb]->[83.2mb]/[989.8mb], all_pools {[young] [269.5mb]->[4.1mb]/[273mb]}{[survivor] [34.1mb]->[34.1mb]/[34.1mb]}{[old] [7.2mb]->[44.9mb]/[682.6mb]}
2021-08-16T19:27:04.033614353Z 19:27:04.028 elasticsearch> [19:27:04.027Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][359] overhead, spent [7.5s] collecting in the last [8.3s]
...
2021-08-16T19:28:46.770105932Z 19:28:46.768 elasticsearch> [19:28:46.767Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][young][459][5] duration [2.2s], collections [1]/[2.7s], total [2.2s]/[9.9s], memory [350.3mb]->[77.9mb]/[989.8mb], all_pools {[young] [271.2mb]->[1.8mb]/[273mb]}{[survivor] [34.1mb]->[13.8mb]/[34.1mb]}{[old] [44.9mb]->[63mb]/[682.6mb]}
2021-08-16T19:28:46.773394585Z 19:28:46.771 elasticsearch> [19:28:46.768Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][459] overhead, spent [2.2s] collecting in the last [2.7s]

我只是想知道是否有什么(除了重试失败)我可以调整或调整来解决这些问题。

我的一些测试执行以下“删除所有”查询以从所有索引中删除所有文档:

RestHighLevelClient client = new RestHighLevelClient(RestClient.builder(new HttpHost(
            elasticHost, Integer.parseInt(elasticPort))));
client.indices().delete(new DeleteIndexRequest("_all"), RequestOptions.DEFAULT);

我想知道这个请求是否会导致 Elasticsearch 运行长时间/完整的垃圾收集?请注意,我的测试套件非常小,它可能会创建 10-15 个索引,并且这些索引中的索引不超过 100 个文档,即我不会创建数百或数千个文档。

以下是 Elasticsearch 状态的一些详细信息。请注意,这是在测试套件完成之后,所以一些测试将执行一些清理(一些索引将被删除等),所以其中一些可能没有用。如果我可以提供更多信息,请告诉我。 请注意,这些统计信息来自测试套件的成功运行,我还没有失败运行的统计信息。

_cluster/stats?pretty&human

{
  "_nodes" : {
    "total" : 1,
    "successful" : 1,
    "failed" : 0
  },
  "cluster_name" : "docker-cluster",
  "cluster_uuid" : "jYnioXskQvi23hDVekWGeA",
  "timestamp" : 1635420632189,
  "status" : "green",
  "indices" : {
    "count" : 16,
    "shards" : {
      "total" : 32,
      "primaries" : 32,
      "replication" : 0.0,
      "index" : {
        "shards" : {
          "min" : 2,
          "max" : 2,
          "avg" : 2.0
        },
        "primaries" : {
          "min" : 2,
          "max" : 2,
          "avg" : 2.0
        },
        "replication" : {
          "min" : 0.0,
          "max" : 0.0,
          "avg" : 0.0
        }
      }
    },
    "docs" : {
      "count" : 50,
      "deleted" : 8
    },
    "store" : {
      "size" : "436.5kb",
      "size_in_bytes" : 447012,
      "reserved" : "0b",
      "reserved_in_bytes" : 0
    },
    "fielddata" : {
      "memory_size" : "7.4kb",
      "memory_size_in_bytes" : 7592,
      "evictions" : 0
    },
    "query_cache" : {
      "memory_size" : "0b",
      "memory_size_in_bytes" : 0,
      "total_count" : 0,
      "hit_count" : 0,
      "miss_count" : 0,
      "cache_size" : 0,
      "cache_count" : 0,
      "evictions" : 0
    },
    "completion" : {
      "size" : "0b",
      "size_in_bytes" : 0
    },
    "segments" : {
      "count" : 40,
      "memory" : "157.4kb",
      "memory_in_bytes" : 161232,
      "terms_memory" : "119.9kb",
      "terms_memory_in_bytes" : 122848,
      "stored_fields_memory" : "19kb",
      "stored_fields_memory_in_bytes" : 19520,
      "term_vectors_memory" : "0b",
      "term_vectors_memory_in_bytes" : 0,
      "norms_memory" : "4.1kb",
      "norms_memory_in_bytes" : 4288,
      "points_memory" : "0b",
      "points_memory_in_bytes" : 0,
      "doc_values_memory" : "14.2kb",
      "doc_values_memory_in_bytes" : 14576,
      "index_writer_memory" : "0b",
      "index_writer_memory_in_bytes" : 0,
      "version_map_memory" : "0b",
      "version_map_memory_in_bytes" : 0,
      "fixed_bit_set" : "0b",
      "fixed_bit_set_memory_in_bytes" : 0,
      "max_unsafe_auto_id_timestamp" : -1,
      "file_sizes" : { }
    },
    "mappings" : {
      "field_types" : [
        {
          "name" : "boolean",
          "count" : 9,
          "index_count" : 2
        },
        {
          "name" : "date",
          "count" : 22,
          "index_count" : 1
        },
        {
          "name" : "double",
          "count" : 2,
          "index_count" : 1
        },
        {
          "name" : "integer",
          "count" : 10,
          "index_count" : 1
        },
        {
          "name" : "join",
          "count" : 1,
          "index_count" : 1
        },
        {
          "name" : "keyword",
          "count" : 169,
          "index_count" : 9
        },
        {
          "name" : "long",
          "count" : 20,
          "index_count" : 6
        },
        {
          "name" : "object",
          "count" : 23,
          "index_count" : 5
        },
        {
          "name" : "text",
          "count" : 61,
          "index_count" : 6
        }
      ]
    },
    "analysis" : {
      "char_filter_types" : [
        {
          "name" : "html_strip",
          "count" : 1,
          "index_count" : 1
        }
      ],
      "tokenizer_types" : [
        {
          "name" : "path_hierarchy",
          "count" : 1,
          "index_count" : 1
        }
      ],
      "filter_types" : [
        {
          "name" : "length",
          "count" : 1,
          "index_count" : 1
        },
        {
          "name" : "limit",
          "count" : 1,
          "index_count" : 1
        },
        {
          "name" : "pattern_capture",
          "count" : 1,
          "index_count" : 1
        }
      ],
      "analyzer_types" : [
        {
          "name" : "custom",
          "count" : 2,
          "index_count" : 1
        },
        {
          "name" : "pattern",
          "count" : 1,
          "index_count" : 1
        },
        {
          "name" : "standard",
          "count" : 1,
          "index_count" : 1
        }
      ],
      "built_in_char_filters" : [ ],
      "built_in_tokenizers" : [
        {
          "name" : "icu_tokenizer",
          "count" : 1,
          "index_count" : 1
        }
      ],
      "built_in_filters" : [
        {
          "name" : "icu_normalizer",
          "count" : 2,
          "index_count" : 1
        }
      ],
      "built_in_analyzers" : [ ]
    }
  },
  "nodes" : {
    "count" : {
      "total" : 1,
      "coordinating_only" : 0,
      "data" : 1,
      "ingest" : 1,
      "master" : 1,
      "remote_cluster_client" : 1
    },
    "versions" : [
      "7.10.2"
    ],
    "os" : {
      "available_processors" : 12,
      "allocated_processors" : 12,
      "names" : [
        {
          "name" : "Linux",
          "count" : 1
        }
      ],
      "pretty_names" : [
        {
          "pretty_name" : "openSUSE Leap 15.3",
          "count" : 1
        }
      ],
      "mem" : {
        "total" : "24.9gb",
        "total_in_bytes" : 26751438848,
        "free" : "17.3gb",
        "free_in_bytes" : 18664411136,
        "used" : "7.5gb",
        "used_in_bytes" : 8087027712,
        "free_percent" : 70,
        "used_percent" : 30
      }
    },
    "process" : {
      "cpu" : {
        "percent" : 0
      },
      "open_file_descriptors" : {
        "min" : 275,
        "max" : 275,
        "avg" : 275
      }
    },
    "jvm" : {
      "max_uptime" : "6.6m",
      "max_uptime_in_millis" : 396405,
      "versions" : [
        {
          "version" : "11.0.12",
          "vm_name" : "OpenJDK 64-Bit Server VM",
          "vm_version" : "11.0.12+7-suse-3.59.1-x8664",
          "vm_vendor" : "Oracle Corporation",
          "bundled_jdk" : true,
          "using_bundled_jdk" : false,
          "count" : 1
        }
      ],
      "mem" : {
        "heap_used" : "292.8mb",
        "heap_used_in_bytes" : 307096056,
        "heap_max" : "989.8mb",
        "heap_max_in_bytes" : 1037959168
      },
      "threads" : 68
    },
    "fs" : {
      "total" : "250.9gb",
      "total_in_bytes" : 269490393088,
      "free" : "243.9gb",
      "free_in_bytes" : 261967937536,
      "available" : "231.1gb",
      "available_in_bytes" : 248207265792
    },
    "plugins" : [
      {
        "name" : "analysis-icu",
        "version" : "7.10.2",
        "elasticsearch_version" : "7.10.2",
        "java_version" : "1.8",
        "description" : "The ICU Analysis plugin integrates the Lucene ICU module into Elasticsearch, adding ICU-related analysis components.",
        "classname" : "org.elasticsearch.plugin.analysis.icu.AnalysisICUPlugin",
        "extended_plugins" : [ ],
        "has_native_controller" : false
      }
    ],
    "network_types" : {
      "transport_types" : {
        "netty4" : 1
      },
      "http_types" : {
        "netty4" : 1
      }
    },
    "discovery_types" : {
      "single-node" : 1
    },
    "packaging_types" : [
      {
        "flavor" : "oss",
        "type" : "docker",
        "count" : 1
      }
    ],
    "ingest" : {
      "number_of_pipelines" : 0,
      "processor_stats" : { }
    }
  }
}

_nodes/stats?pretty&human

太大无法粘贴,复制到这里:https://pastebin.com/EBghcNuQ

这是上面pastebin的摘录,看起来年轻一代几乎已经用尽了,但是,即便如此,如果需要这么长时间(距离这篇文章顶部的警告7秒) ) 收集?

"pools" : {
    "young" : {
      "used" : "219.7mb",
      "used_in_bytes" : 230396648,
      "max" : "273mb",
      "max_in_bytes" : 286326784,
      "peak_used" : "273mb",
      "peak_used_in_bytes" : 286326784,
      "peak_max" : "273mb",
      "peak_max_in_bytes" : 286326784
    },
    "survivor" : {
      "used" : "9.9mb",
      "used_in_bytes" : 10391856,
      "max" : "34.1mb",
      "max_in_bytes" : 35782656,
      "peak_used" : "34.1mb",
      "peak_used_in_bytes" : 35782656,
      "peak_max" : "34.1mb",
      "peak_max_in_bytes" : 35782656
    },
    "old" : {
      "used" : "63.2mb",
      "used_in_bytes" : 66307552,
      "max" : "682.6mb",
      "max_in_bytes" : 715849728,
      "peak_used" : "63.2mb",
      "peak_used_in_bytes" : 66307552,
      "peak_max" : "682.6mb",
      "peak_max_in_bytes" : 715849728
    }
  }
},
"threads" : {
  "count" : 68,
  "peak_count" : 68
},
"gc" : {
  "collectors" : {
    "young" : {
      "collection_count" : 10,
      "collection_time" : "115ms",
      "collection_time_in_millis" : 115
    },
    "old" : {
      "collection_count" : 2,
      "collection_time" : "73ms",
      "collection_time_in_millis" : 73
    }
  }
},

_cluster/settings?pretty

{
  "persistent" : { },
  "transient" : { }
}

_cat/nodes?h=heap*&v

heap.current heap.percent heap.max
     117.5mb           11  989.8mb

_cat/nodes?v

ip         heap.percent ram.percent cpu load_1m load_5m load_15m node.role master name
172.17.0.5           12          30   0    0.07    0.05     0.09 dimr      *      aabedf9ca3ca

_cat/fielddata?v

id                     host       ip         node         field size
_eJ-A7pFSoKMfuDKs63SLg 172.17.0.5 172.17.0.5 aabedf9ca3ca _id    1kb

_cat/shards

workspace-archive-cleanup-job_item-000001 1 p STARTED 1  4.1kb 172.17.0.5 aabedf9ca3ca
workspace-archive-cleanup-job_item-000001 0 p STARTED 5   19kb 172.17.0.5 aabedf9ca3ca
assigndeviceids_item-000001               1 p STARTED 0   208b 172.17.0.5 aabedf9ca3ca
assigndeviceids_item-000001               0 p STARTED 4 29.3kb 172.17.0.5 aabedf9ca3ca

_cat/allocation?v

shards disk.indices disk.used disk.avail disk.total disk.percent host       ip         node
     4       52.7kb    19.7gb    231.2gb    250.9gb            7 172.17.0.5 172.17.0.5 aabedf9ca3ca

JVM 选项:

elastic+     7  0.5  5.2 7594004 1360460 ?     Sl   06:14   1:00 /usr/lib64/jvm/java-11-openjdk-11/bin/java
-Xshare:auto
-Des.networkaddress.cache.ttl=60
-Des.networkaddress.cache.negative.ttl=10
-XX:+AlwaysPreTouch
-Xss1m
-Djava.awt.headless=true
-Dfile.encoding=UTF-8
-Djna.nosys=true
-XX:-OmitStackTraceInFastThrow
-Dio.netty.noUnsafe=true
-Dio.netty.noKeySetOptimization=true
-Dio.netty.recycler.maxCapacityPerThread=0
-Dio.netty.allocator.numDirectArenas=0
-Dlog4j.shutdownHookEnabled=false
-Dlog4j2.disable.jmx=true
-Djava.locale.providers=SPI,COMPAT
-Xms1g
-Xmx1g
-XX:+UseConcMarkSweepGC
-XX:CMSInitiatingOccupancyFraction=75
-XX:+UseCMSInitiatingOccupancyOnly
-Djava.io.tmpdir=/tmp/elasticsearch-9557497152372782786
-XX:+HeapDumpOnOutOfMemoryError
-XX:HeapDumpPath=data
-XX:ErrorFile=logs/hs_err_pid%p.log -Xlog:gc*,gc+age=trace,safepoint:file=logs/gc.log:utctime,pid,tags:filecount=32,filesize=64m -Des.cgroups.hierarchy.override=/
-Dcom.sun.management.jmxremote
-Dcom.sun.management.jmxremote.port=9010
-Dcom.sun.management.jmxremote.rmi.port=9010
-Djava.rmi.server.hostname=127.0.0.1
-Dcom.sun.management.jmxremote.local.only=false
-Dcom.sun.management.jmxremote.authenticate=false
-Dcom.sun.management.jmxremote.ssl=false
-XX:MaxDirectMemorySize=536870912
-Des.path.home=/usr/share/elasticsearch
-Des.path.conf=/usr/share/elasticsearch/config
-Des.distribution.flavor=oss
-Des.distribution.type=docker
-Des.bundled_jdk=true 
-cp /usr/share/elasticsearch/lib/* org.elasticsearch.bootstrap.Elasticsearch -Ediscovery.type=single-node

从日志中提取直到第一次 GC 警告:

elasticsearch> OpenJDK 64-Bit Server VM warning: Option UseConcMarkSweepGC was deprecated in version 9.0 and will likely be removed in a future release.
elasticsearch> [19:20:31.086Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] version[7.10.2], pid[8], build[oss/docker/747e1cc71def077253878a59143c1f785afa92b9/2021-01-13T00:42:12.435326Z], OS[Linux/3.10.0-862.6.3.el7.x86_64/amd64], JVM[Oracle Corporation/OpenJDK 64-Bit Server VM/11.0.11/11.0.11+9-suse-lp152.2.12.1-x8664]
elasticsearch> [19:20:31.097Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] JVM home [/usr/lib64/jvm/java-11-openjdk-11], using bundled JDK [false]
elasticsearch> [19:20:31.097Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] JVM arguments [-Xshare:auto, -Des.networkaddress.cache.ttl=60, -Des.networkaddress.cache.negative.ttl=10, -XX:+AlwaysPreTouch, -Xss1m, -Djava.awt.headless=true, -Dfile.encoding=UTF-8, -Djna.nosys=true, -XX:-OmitStackTraceInFastThrow, -Dio.netty.noUnsafe=true, -Dio.netty.noKeySetOptimization=true, -Dio.netty.recycler.maxCapacityPerThread=0, -Dio.netty.allocator.numDirectArenas=0, -Dlog4j.shutdownHookEnabled=false, -Dlog4j2.disable.jmx=true, -Djava.locale.providers=SPI,COMPAT, -Xms1g, -Xmx1g, -XX:+UseConcMarkSweepGC, -XX:CMSInitiatingOccupancyFraction=75, -XX:+UseCMSInitiatingOccupancyOnly, -Djava.io.tmpdir=/tmp/elasticsearch-11638220902719581930, -XX:+HeapDumpOnOutOfMemoryError, -XX:HeapDumpPath=data, -XX:ErrorFile=logs/hs_err_pid%p.log, -Xlog:gc*,gc+age=trace,safepoint:file=logs/gc.log:utctime,pid,tags:filecount=32,filesize=64m, -Des.cgroups.hierarchy.override=/, -Dcom.sun.management.jmxremote, -Dcom.sun.management.jmxremote.port=9010, -Dcom.sun.management.jmxremote.rmi.port=9010, -Djava.rmi.server.hostname=127.0.0.1, -Dcom.sun.management.jmxremote.local.only=false, -Dcom.sun.management.jmxremote.authenticate=false, -Dcom.sun.management.jmxremote.ssl=false, -XX:MaxDirectMemorySize=536870912, -Des.path.home=/usr/share/elasticsearch, -Des.path.conf=/usr/share/elasticsearch/config, -Des.distribution.flavor=oss, -Des.distribution.type=docker, -Des.bundled_jdk=true]
elasticsearch> [19:20:43.124Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [aggs-matrix-stats]
elasticsearch> [19:20:43.125Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [analysis-common]
elasticsearch> [19:20:43.125Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [geo]
elasticsearch> [19:20:43.125Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [ingest-common]
elasticsearch> [19:20:43.126Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [ingest-geoip]
elasticsearch> [19:20:43.126Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [ingest-user-agent]
elasticsearch> [19:20:43.126Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [kibana]
elasticsearch> [19:20:43.127Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [lang-expression]
elasticsearch> [19:20:43.127Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [lang-mustache]
elasticsearch> [19:20:43.127Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [lang-painless]
elasticsearch> [19:20:43.127Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [mapper-extras]
elasticsearch> [19:20:43.128Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [parent-join]
elasticsearch> [19:20:43.128Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [percolator]
elasticsearch> [19:20:43.128Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [rank-eval]
elasticsearch> [19:20:43.129Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [reindex]
elasticsearch> [19:20:43.129Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [repository-url]
elasticsearch> [19:20:43.129Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded module [transport-netty4]
elasticsearch> [19:20:43.130Z #205.001 INFO  -            -   ] org.elasticsearch.plugins.PluginsService: [a65fc811a8a8] loaded plugin [analysis-icu]
elasticsearch> [19:20:43.761Z #205.001 INFO  -            -   ] org.elasticsearch.env.NodeEnvironment: [a65fc811a8a8] using [1] data paths, mounts [[/ (rootfs)]], net usable_space [9.4gb], net total_space [9.9gb], types [rootfs]
elasticsearch> [19:20:43.762Z #205.001 INFO  -            -   ] org.elasticsearch.env.NodeEnvironment: [a65fc811a8a8] heap size [989.8mb], compressed ordinary object pointers [true]
elasticsearch> [19:20:43.962Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] node name [a65fc811a8a8], node ID [zinhYoeoQ_q-fGulfFN42g], cluster name [docker-cluster], roles [master, remote_cluster_client, data, ingest]
elasticsearch> [19:20:53.399Z #205.001 INFO  -            -   ] org.elasticsearch.transport.NettyAllocator: [a65fc811a8a8] creating NettyAllocator with the following configs: [name=unpooled, suggested_max_allocation_size=1mb, factors={es.unsafe.use_unpooled_allocator=null, g1gc_enabled=false, g1gc_region_size=0b, heap_size=989.8mb}]
elasticsearch> [19:20:53.580Z #205.001 INFO  -            -   ] org.elasticsearch.discovery.DiscoveryModule: [a65fc811a8a8] using discovery type [single-node] and seed hosts providers [settings]
elasticsearch> [19:20:54.126Z #205.001 WARN  -            -   ] org.elasticsearch.gateway.DanglingIndicesState: [a65fc811a8a8] gateway.auto_import_dangling_indices is disabled, dangling indices will not be automatically detected or imported and must be managed manually
elasticsearch> [19:20:54.356Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] initialized
elasticsearch> [19:20:54.357Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] starting ...
elasticsearch> [19:20:54.614Z #205.001 INFO  -            -   ] org.elasticsearch.transport.TransportService: [a65fc811a8a8] publish_address {172.17.0.21:9300}, bound_addresses {0.0.0.0:9300}
elasticsearch> [19:20:55.452Z #205.001 INFO  -            -   ] org.elasticsearch.cluster.coordination.Coordinator: [a65fc811a8a8] setting initial configuration to VotingConfiguration{zinhYoeoQ_q-fGulfFN42g}
elasticsearch> [19:20:55.729Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.service.MasterService: [a65fc811a8a8] elected-as-master ([1] nodes joined)[{a65fc811a8a8}{zinhYoeoQ_q-fGulfFN42g}{p4QxXvFvQgWh3Pi-VTmk4g}{172.17.0.21}{172.17.0.21:9300}{dimr} elect leader, _BECOME_MASTER_TASK_, _FINISH_ELECTION_], term: 1, version: 1, delta: master node changed {previous [], current [{a65fc811a8a8}{zinhYoeoQ_q-fGulfFN42g}{p4QxXvFvQgWh3Pi-VTmk4g}{172.17.0.21}{172.17.0.21:9300}{dimr}]}
elasticsearch> [19:20:55.809Z #205.027 INFO  -            -   ] org.elasticsearch.cluster.coordination.CoordinationState: [a65fc811a8a8] cluster UUID set to [FDeYMVECTAGBjQSMxOtUGw]
elasticsearch> [19:20:55.859Z #205.024 INFO  -            -   ] org.elasticsearch.cluster.service.ClusterApplierService: [a65fc811a8a8] master node changed {previous [], current [{a65fc811a8a8}{zinhYoeoQ_q-fGulfFN42g}{p4QxXvFvQgWh3Pi-VTmk4g}{172.17.0.21}{172.17.0.21:9300}{dimr}]}, term: 1, version: 1, reason: Publication{term=1, version=1}
elasticsearch> [19:20:55.913Z #205.001 INFO  -            -   ] org.elasticsearch.http.AbstractHttpServerTransport: [a65fc811a8a8] publish_address {172.17.0.21:9200}, bound_addresses {0.0.0.0:9200}
elasticsearch> [19:20:55.913Z #205.001 INFO  -            -   ] org.elasticsearch.node.Node: [a65fc811a8a8] started
elasticsearch> [19:20:55.955Z #205.028 INFO  -            -   ] org.elasticsearch.gateway.GatewayService: [a65fc811a8a8] recovered [0] indices into cluster_state
elasticsearch> [19:26:02.442Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.metadata.MetadataCreateIndexService: [a65fc811a8a8] [hudzh_item-000001] creating index, cause [api], templates [], shards [2]/[0]
elasticsearch> [19:26:03.639Z #205.028 WARN  -            -   ] org.elasticsearch.cluster.service.MasterService: [a65fc811a8a8] took [17.5s], which is over [10s], to compute cluster state update for [create-index [hudzh_item-000001], cause [api]]
elasticsearch> [19:26:15.364Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.metadata.MetadataCreateIndexService: [a65fc811a8a8] [assigndeviceids_item-000001] creating index, cause [api], templates [], shards [2]/[0]
elasticsearch> [19:26:19.387Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.routing.allocation.AllocationService: [a65fc811a8a8] Cluster health status changed from [YELLOW] to [GREEN] (reason: [shards started [[assigndeviceids_item-000001][0], [assigndeviceids_item-000001][1]]]).
elasticsearch> [19:26:20.559Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.metadata.MetadataMappingService: [a65fc811a8a8] [assigndeviceids_item-000001/-yvbBBiRR_maejjkGxNNlA] create_mapping [_doc]
elasticsearch> [19:26:27.288Z #205.028 INFO  -            -   ] org.elasticsearch.cluster.metadata.MetadataMappingService: [a65fc811a8a8] [assigndeviceids_item-000001/-yvbBBiRR_maejjkGxNNlA] update_mapping [_doc]
elasticsearch> [19:27:04.025Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][young][359][4] duration [7.5s], collections [1]/[8.3s], total [7.5s]/[7.7s], memory [310.9mb]->[83.2mb]/[989.8mb], all_pools {[young] [269.5mb]->[4.1mb]/[273mb]}{[survivor] [34.1mb]->[34.1mb]/[34.1mb]}{[old] [7.2mb]->[44.9mb]/[682.6mb]}
elasticsearch> [19:27:04.027Z #205.021 WARN  -            -   ] org.elasticsearch.monitor.jvm.JvmGcMonitorService: [a65fc811a8a8] [gc][359] overhead, spent [7.5s] collecting in the last [8.3s]

注意上面日志的第一行:

elasticsearch> OpenJDK 64-Bit Server VM warning: Option UseConcMarkSweepGC was deprecated in version 9.0 and will likely be removed in a future release.

可能相关:

https://github.com/elastic/elasticsearch/issues/43911 https://github.com/elastic/elasticsearch/issues/36828

请注意,我对失败的测试环境和我没有发现任何问题的其他环境进行了比较。

所有环境都使用相同的默认 CMS 垃圾收集器。在我的其他环境中,虽然 Elasticsearch 分配了 8GB,但在我的测试环境中分配了默认的 1GB。

我有点认为 1GB 就足够了,因为如上所述,关于创建的索引和文档的数量,我对 Elasticsearch 的测试并不是特别繁重?

更新:

我尝试将内存从 1GB 增加到 2GB,但是,我仍然看到 GC 警告。

我还在每个测试类中添加了一些日志,以便在每个测试开始时打印出一些统计数据并进行一些比较。

在测试失败并且我看到 GC 警告的运行期间,我发现在一个测试开始和另一个测试开始之间有 3 个年轻的垃圾回收,总共花费了 6 秒,而不是通常的 50 毫秒左右:

但是,这仅发生在 5 次测试运行中的 1 次中,这会导致一个或多个测试失败。通常在大多数情况下,每个年轻的 GC 大约需要 50 毫秒或更短的时间。在通过和失败的情况下,测试的运行顺序是相同的。

有趣的是,这似乎与某些测试运行的“全部删除”操作相吻合,这是同一日志的另一部分:

可能是删除 78 个文档引发了一些年轻的 GC?我只是不确定仅在某些情况下会发生什么?

来自同一日志的其他部分:

这里还有一部分集群统计信息。我注意到一个特定的测试正在创建大量索引 (24) 和相关的分片。随之而来的下一个测试执行“全部删除”请求。也许这可能是 GC 警告的原因:

final DeleteIndexRequest deleteAll = new DeleteIndexRequest("_all");
client.indices().delete(deleteAll, RequestOptions.DEFAULT);

【问题讨论】:

  • 我可能会建议您将堆大小增加到 2GB,然后使用更新版本的 Elasticsearch(最新版本是 7.15)
  • “我只是想知道是否有任何东西(除了重试失败)我可以调整或调整以解决这些问题” - 摆脱 CMS,它在 jdk-14 中被删除。增加堆,便宜,简单,适用于绝大多数情况。更长一点:即使对于“不太大的应用程序”,您也会惊讶地发现需要多少堆,GC 拥有的余量越大,它的性能就越好。

标签: java docker elasticsearch garbage-collection


【解决方案1】:

我将从这个问题开始: 这是 Elasticsearch、JAVA 或您的网络环境(包括您的测试代码)的问题吗?

为了尝试确定我会提出以下步骤:

  1. 使用 Elasticsearch 查看您为 JVM 运行的内容。根据您的日志,您正在使用:
  • -XX:+UseConcMarkSweepGC
  • -XX:CMSInitiatingOccupancyFracion=75I
  • -XX:+UseCMSInitiatingOccupancyOnly

对于你提到的这个问题:

elasticsearch> OpenJDK 64 位服务器虚拟机警告:选项 UseConcMarkSweepGC 在 9.0 版中已弃用,可能会 在未来的版本中删除。

您可以尝试其中的每一种,看看它是否能解决您的问题:

  • 切换到 G1 GC 算法。它在 JAVA 9 之后成为默认设置。与 CMS 相比,它更容易调整,并且具有更少的参数。
  • 切换到 Z GC 算法。适用于您的版本 JAVA 11。其目标是将 GC 暂停时间保持在 10 毫秒以下。

不保证,但我认为尝试这样做应该可以消除您的 JVM 配置作为根本原因。

  1. 运行测试后运行 Visual VM 并执行堆转储和/或线程转储。您的测试中是否有可能导致堆或线程问题?这可以告诉你。 Visual VM在你的版本中不再打包JAVA,但可以从Visual VM下载。

  2. 运行 GCEasy 以查看它显示的分析。
    它有一个有限的免费版本,具有以下功能:

  • 5 次上传/用户/月
  • 5 次 API 调用/用户/月
  • 10 mb 文件大小/上传
  • 在我们的云中运行
  • 数据存储在我们的云中
  • 标准支持
  • 几张图

分析 gc 日志是他们的专长。我认为使用免费版本会指出您的问题或提供提示,您可以使用我在此处列出的其他步骤中的信息来扩展信息。

  1. 通常我会同意 1GB 的内存就足够了,但我会将它提高到 2GB,因为这是我们在生产环境中运行的,因为它可以解决问题。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2011-01-21
    • 1970-01-01
    • 2018-12-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2010-12-13
    • 2012-05-04
    相关资源
    最近更新 更多