【发布时间】:2020-03-03 08:16:41
【问题描述】:
编辑:简化我的问题
我将事件日志流式传输到单个弹性搜索索引“日志”中。
这些事件日志中的每一个都代表某个长时间运行的进程中的单个事件。事件日志的状态为 queued->init->running->finished。
现在我想查询我的索引以查找正在运行的作业...但是如果我搜索 status="running",我还将获得稍后已完成的作业。作为一个人类,我会寻找带有“正在运行”事件但没有“已完成”事件的日志。
假设我必须耙树叶,打扫房间,然后吃午饭。我的事件日志如下所示: 即
{job: "rake leaves" status: "queued"}
{job: "clean house" status: "queued"}
{job: "eat lunch" status: "queued"}
{job: "rake leaves" status: "started"}
{job: "rake leaves" status: "running"}
{job: "rake leaves" status: "finished"}
{job: "clean house" status: "started"}
{job: "clean house" status: "finished"}
{job: "clean house" status: "running"}
{job: "eat lunch" status: "started"}
{job: "eat lunch" status: "running"}
如何找到尚未完成的正在运行的作业?在这种情况下,吃午饭 是唯一正在运行的作业。我将扩展它以查找尚未开始的排队作业。状态无关紧要。
我目前的思路是使用反向嵌套聚合来冒泡所有状态,然后从那里过滤掉我不想要的项目。我也可能会滥用带有 min_doc_count 的术语聚合器来获得我需要的东西。
例子
curl -H "Content-Type: application/json" -XPUT "http://localhost:9200/jobs/" -d'
{
"mappings": {
"event": {
"properties": {
"name": {
"type": "keyword"
},
"status": {
"type": "keyword"
}
}
}
}
}'
- 一些示例数据
curl -H "Content-Type: application/json" -XPOST "http://localhost:9200/jobs/_bulk" -d'
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job0", "status":"init"}
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job1", "status":"init"}
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job2", "status":"init"}
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job0", "status":"running"}
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job1", "status":"running"}
{"index":{"_index":"jobs","_type":"event"}}
{"name":"job0", "status":"finished"}
'
- 示例查询
curl -H "Content-Type: application/json" -XGET http://localhost:9200/logs/event/_search -d'
{
"aggs": {
"duplicateNames": {
"terms": {
"field": "name",
"min_doc_count": 2
}
}
}
}' | python -m json.tool
【问题讨论】:
标签: elasticsearch elasticsearch-aggregation