【问题标题】:Elasticsearch, find documents without matching documents containing valueElasticsearch,查找不匹配包含值的文档的文档
【发布时间】:2020-03-03 08:16:41
【问题描述】:

编辑:简化我的问题

我将事件日志流式传输到单个弹性搜索索引“日志”中。

这些事件日志中的每一个都代表某个长时间运行的进程中的单个事件。事件日志的状态为 queued->init->running->finished。

现在我想查询我的索引以查找正在运行的作业...但是如果我搜索 status="running",我还将获得稍后已完成的作业。作为一个人类,我会寻找带有“正在运行”事件但没有“已完成”事件的日志。

假设我必须耙树叶,打扫房间,然后吃午饭。我的事件日志如下所示: 即

{job: "rake leaves" status: "queued"}
{job: "clean house" status: "queued"}
{job: "eat lunch" status: "queued"}
{job: "rake leaves" status: "started"}
{job: "rake leaves" status: "running"}
{job: "rake leaves" status: "finished"}
{job: "clean house" status: "started"}
{job: "clean house" status: "finished"}
{job: "clean house" status: "running"}
{job: "eat lunch" status: "started"}
{job: "eat lunch" status: "running"}

如何找到尚未完成的正在运行的作业?在这种情况下,吃午饭 是唯一正在运行的作业。我将扩展它以查找尚未开始的排队作业。状态无关紧要。

我目前的思路是使用反向嵌套聚合来冒泡所有状态,然后从那里过滤掉我不想要的项目。我也可能会滥用带有 min_doc_count 的术语聚合器来获得我需要的东西。

例子

    curl -H "Content-Type: application/json" -XPUT "http://localhost:9200/jobs/" -d'
    {
       "mappings": {
          "event": {
             "properties": {
                "name": {
                   "type": "keyword"
                },
                "status": {
                   "type": "keyword"
                }
             }
          }
       }
    }'
  • 一些示例数据

    curl -H "Content-Type: application/json" -XPOST "http://localhost:9200/jobs/_bulk" -d'
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job0", "status":"init"}
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job1", "status":"init"}
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job2", "status":"init"}
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job0", "status":"running"}
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job1", "status":"running"}
    {"index":{"_index":"jobs","_type":"event"}}
    {"name":"job0", "status":"finished"}
    '
  • 示例查询
    curl -H "Content-Type: application/json" -XGET http://localhost:9200/logs/event/_search -d'
    {
        "aggs": {
            "duplicateNames": {
                "terms": {
                    "field": "name",
                    "min_doc_count": 2
                }
            }
        }
    }' | python -m json.tool

【问题讨论】:

    标签: elasticsearch elasticsearch-aggregation


    【解决方案1】:

    所以,说实话!我确实了解上下文和所有内容。但是,我不太明白你的问题。你想在这里实现什么目标?

    也许在你(做得很好的解释:D)的最后把问题清楚地放在最后。 :)

    【讨论】:

    • 您好,我重新格式化了问题。我有与作业名称和作业状态一起流式传输的日志数据。我想搜索正在运行但尚未完成的作业
    猜你喜欢
    • 1970-01-01
    • 2017-11-19
    • 2014-08-24
    • 2021-12-06
    • 1970-01-01
    • 2014-12-16
    • 2015-10-31
    • 2021-08-24
    • 1970-01-01
    相关资源
    最近更新 更多