【问题标题】:How to Query elasticsearch index with nested and non nested fields如何使用嵌套和非嵌套字段查询弹性搜索索引
【发布时间】:2020-02-25 03:23:09
【问题描述】:

我有一个具有以下映射的弹性搜索索引:

PUT /student_detail
{
    "mappings" : {
        "properties" : {
            "id" : { "type" : "long" },
            "name" : { "type" : "text" },
            "email" : { "type" : "text" },
            "age" : { "type" : "text" },
            "status" : { "type" : "text" },
            "tests":{ "type" : "nested" }
        }
    }
}

存储的数据格式如下:

{
  "id": 123,
  "name": "Schwarb",
  "email": "abc@gmail.com",
  "status": "current",
  "age": 14,
  "tests": [
    {
      "test_id": 587,
      "test_score": 10
    },
    {
      "test_id": 588,
      "test_score": 6
    }
  ]
}

我希望能够查询学生的姓名,如 '%warb%' 和电子邮件,如 '%gmail.com%' 和 id 587 的测试得分 > 5 等。所需的高水平可以是把类似下面的东西,不知道实际的查询是什么,为下面这个混乱的查询道歉

GET developer_search/_search
{
  "query": {
    "bool": {
      "must": [
        {
          "match": {
            "name": "abc"
          }
        },
        {
          "nested": {
            "path": "tests",
            "query": {
              "bool": {
                "must": [
                  {
                    "term": {
                      "tests.test_id": IN [587]
                    }
                  },
                   {
                    "term": {
                      "tests.test_score": >= some value
                    }
                  }
                ]
              }
            }
          }
        }
      ]
    }
  }
}

查询必须灵活,以便我们可以输入动态测试 ID 及其各自的分数过滤器以及年龄、姓名、状态等嵌套字段之外的字段

【问题讨论】:

    标签: elasticsearch elastic-stack


    【解决方案1】:

    类似的东西?

    GET student_detail/_search
    {
      "query": {
        "bool": {
          "must": [
            {
              "wildcard": {
                "name": {
                  "value": "*warb*"
                }
              }
            },
            {
              "wildcard": {
                "email": {
                  "value": "*gmail.com*"
                }
              }
            },
            {
              "nested": {
                "path": "tests",
                "query": {
                  "bool": {
                    "must": [
                      {
                        "term": {
                          "tests.test_id": 587
                        }
                      },
                      {
                        "range": {
                          "tests.test_score": {
                            "gte": 5
                          }
                        }
                      }
                    ]
                  }
                },
                "inner_hits": {}
              }
            }
          ]
        }
      }
    }
    

    你正在寻找内在的命中。

    【讨论】:

      【解决方案2】:

      您必须使用Ngram Tokenizer,因为出于性能原因不得使用通配符搜索,我不建议使用它。

      将您的映射更改为以下,您可以创建自己的Analyzer,我在下面的映射中完成了。

      elasticsearch (albiet lucene) 如何索引语句,首先它将语句或段落分解为单词或标记,然后在该特定字段的倒排索引中索引这些单词。此过程称为Analysis,并且仅适用于text 数据类型。

      因此,现在您只有在这些标记在倒排索引中可用时才能获取文档。

      默认情况下,将应用standard analyzer。我所做的是创建了自己的分析器并使用了 Ngram Tokenizer,它可以创建更多的标记,而不仅仅是单词。

      Life is beautiful 上的默认分析器将是 life、is、beautiful。

      但是使用 Ngram,Life 的标记将是 lif、ife 和 life

      映射:

      PUT student_detail
      {
        "settings": {
          "analysis": {
            "analyzer": {
              "my_analyzer": {
                "tokenizer": "my_tokenizer"
              }
            },
            "tokenizer": {
              "my_tokenizer": {
                "type": "ngram",
                "min_gram": 3,
                "max_gram": 4,
                "token_chars": [
                  "letter",
                  "digit"
                ]
              }
            }
          }
        },
          "mappings" : {
              "properties" : {
                  "id" : { 
                    "type" : "long" 
                  },
                  "name" : { 
                    "type" : "text",
                    "analyzer": "my_analyzer",
                    "fields": {
                      "keyword": {
                        "type": "keyword"
                      }
                    }
                  },
                  "email" : { 
                    "type" : "text",
                    "analyzer": "my_analyzer",
                    "fields": {
                      "keyword": {
                        "type": "keyword"
                      }
                    }
                  },
                  "age" : { 
                    "type" : "text"             <--- I am not sure why this is text. Change it to long or int. Would leave this to you
                  },
                  "status" : { 
                    "type" : "text",
                    "analyzer": "my_analyzer",
                    "fields": {
                      "keyword": {
                        "type": "keyword"
                      }
                    }
                  },
                  "tests":{ 
                    "type" : "nested" 
                  }
              }
          }
      }
      

      请注意,在上面的映射中,我为name、email 和status 创建了一个关键字形式的兄弟字段,如下所示:

      "name":{ 
         "type":"text",
         "analyzer":"my_analyzer",
         "fields":{ 
            "keyword":{ 
               "type":"keyword"
            }
         }
      }
      

      现在您的查询可以像下面这样简单。

      查询:

      POST student_detail/_search
      {
        "query": {
          "bool": {
            "must": [
              {
                "match": {
                  "name": "war"                      <---- Note this. This would even return documents having "Schwarb"
                }
              },
              {
                "match": {
                  "email": "gmail"                   <---- Note this
                }
              },
              {
                "nested": {
                  "path": "tests",
                  "query": {
                    "bool": {
                      "must": [
                        {
                          "term": {
                            "tests.test_id": 587
                          }
                        },
                        {
                          "range": {
                            "tests.test_score": {
                              "gte": 5
                            }
                          }
                        }
                      ]
                    }
                  }
                }
              }
            ]
          }
        }
      }
      

      请注意,对于完全匹配,我将在 keyword 字段中使用 Term Queries,而对于普通搜索或 SQL 中的 LIKE,我将在 text 字段中使用简单的 Match Queries 如果他们使用Ngram Tokenizer。

      另请注意,对于&gt;= 和&lt;=,您需要使用Range Query。

      回应:

      {
        "took" : 233,
        "timed_out" : false,
        "_shards" : {
          "total" : 1,
          "successful" : 1,
          "skipped" : 0,
          "failed" : 0
        },
        "hits" : {
          "total" : {
            "value" : 1,
            "relation" : "eq"
          },
          "max_score" : 3.7260926,
          "hits" : [
            {
              "_index" : "student_detail",
              "_type" : "_doc",
              "_id" : "1",
              "_score" : 3.7260926,
              "_source" : {
                "id" : 123,
                "name" : "Schwarb",
                "email" : "abc@gmail.com",
                "status" : "current",
                "age" : 14,
                "tests" : [
                  {
                    "test_id" : 587,
                    "test_score" : 10
                  },
                  {
                    "test_id" : 588,
                    "test_score" : 6
                  }
                ]
              }
            }
          ]
        }
      }
      

      请注意,当我运行查询时,我会在回复中看到您在问题中提到的文档。

      请阅读我分享的链接。了解这些概念至关重要。希望这会有所帮助!

      【讨论】:

        猜你喜欢
        • 2017-06-11
        • 2021-09-01
        • 2015-09-24
        • 1970-01-01
        • 1970-01-01
        • 2020-06-14
        • 2020-10-12
        • 1970-01-01
        • 2020-07-26
        相关资源
        最近更新 更多