【问题标题】:nGram partial matching & limiting nGram results in multiple field querynGram 部分匹配和限制 nGram 导致多字段查询
【发布时间】:2017-02-16 21:33:26
【问题描述】:

背景:我已通过索引标记化名称(name 字段)以及三元组分析名称(ngram 字段)对名称字段进行了部分搜索。

我已经提升了 name 字段以使精确的标记匹配冒泡到结果的顶部。

问题:我正在尝试实现一个查询,将 nGram 匹配限制为仅匹配查询字符串的某个阈值(例如 80%)的查询。我知道minimum_should_match 似乎是我正在寻找的东西,但我的问题是形成查询以实际产生这些结果。

我的确切标记匹配被提升到顶部,但我仍然得到 每个在 ngram 字段中具有单个匹配三元组的文档。

GIST: Index settings and mapping

索引设置

{
  "my_index": {
    "settings": {
      "index": {
        "number_of_shards": "5",
        "max_result_window": "30000",
        "creation_date": "1475853851937",
        "analysis": {
          "filter": {
            "ngram_filter": {
              "type": "ngram",
              "min_gram": "3",
              "max_gram": "3"
            }
          },
          "analyzer": {
            "ngram_analyzer": {
              "filter": [
                "lowercase",
                "ngram_filter"
              ],
              "type": "custom",
              "tokenizer": "standard"
            }
          }
        },
        "number_of_replicas": "1",
        "uuid": "AuCjcP5sSb-m59bYrprFcw",
        "version": {
          "created": "2030599"
        }
      }
    }
  }
}

索引映射

{
  "my_index": {
    "mappings": {
      "my_type": {
        "properties": {
          "acw": {
            "type": "integer"
          },
          "pcg": {
            "type": "integer"
          },
          "date": {
            "type": "date",
            "format": "strict_date_optional_time||epoch_millis"
          },
          "dob": {
            "type": "date",
            "format": "strict_date_optional_time||epoch_millis"
          },
          "id": {
            "type": "string"
          },
          "name": {
            "type": "string",
            "boost": 10
          },
          "ngram": {
            "type": "string",
            "analyzer": "ngram_analyzer"
          },
          "bdk": {
            "type": "integer"
          },
          "mmw": {
            "type": "integer"
          },
          "mpi": {
            "type": "integer"
          },
          "sex": {
            "type": "string",
            "index": "not_analyzed"
          }
        }
      }
    }
  }
}

解决方案尝试

[GIST:查询尝试] 由于 2 个链接限制而取消链接:( (https://gist.github.com/jordancardwell/2e690013666e7e1da6ef1acee314b4e6)

我尝试了一个多重匹配查询,它给了我正确的搜索结果,但我没有运气省略只匹配单个三元组的名称的结果(比如“odo”三元组里面“ odophilus")

//this matches 'frodo' and sends results to the top, since `name` field is boosted
//  but also matches 'theodore' and 'rodolpho'

{
  "size":100,
  "from":0,
  "query":{
    "multi_match":{
      "query":"frodo",
      "fields":[
        "name",
        "ngram"
      ],
      "type":"best_fields"
    }
  }
}

.

//I then tried to throw in the `minimum_must_match` option
// hoping it would filter out large strings that only had one matching trigram for instance
{
  "size":100,
  "from":0,
  "query":{
    "multi_match":{
      "query":"frodo",
      "fields":[
        "name",
        "ngram"
      ],
      "type":"best_fields",
      "minimum_should_match": "90%",
    }
  }
}

我已经尝试过有意义地手动生成匹配查询,从而允许我仅将 minimum_must_match 应用于 ngram 字段,但似乎无法正确使用语法。

// I then tried to contruct a custom query to just return the `minimum_should_match`d results on the ngram field
// I started with a query produced by using bodybuilder to `and` and `or` my other search criteria together
{
  "query": {
    "bool": {
      "filter": {
        "bool": {
          "must": [
            //each separate field's criteria `must`/`and`ed together
            {
              "query": {
                "bool": {
                  "filter": {
                    "bool": {
                      "should": [
                        //each critereon for a specific field `should`/`or`ed together
                        {
                         //my attempt at getting `ngram` field results.. 
                         // should theoretically only return when field 
                         // contains nothing but matching ngrams 
                         // (i.e. exact matches and other fluke matches)
                          "query": { 
                            "match": {
                              "ngram": {
                                "query": "frodo",
                                "minimum_should_match": "100%"
                              }
                            }
                          }
                        }
                        //... other critereon to be `should`/`or`ed together
                      ]
                    }
                  }
                }
              }
            }
            //... other criteria to be `must`/`and`ed together
          ]
        }
      }
    }
  }
}

谁能看出我做错了什么?

看起来这应该很容易完成,但我肯定遗漏了一些明显的东西。


更新

我使用 _explain=true(使用 sense UI)运行了一个查询,试图了解我的结果。

我在"frod" 的"frod" 和minimum_should_match = 100% 的ngram 字段上查询了match,但我仍然得到与至少一个ngram 匹配的每条记录。 (例如rodolpho,即使它不包含fro)

GIST: test query and results


注意:交叉发布自 [discuss.elastic.co] 稍后会做一个链接,不能发布超过2个:/

(https://discuss.elastic.co/t/ngram-partial-match-limiting-ngram-results-in-multiple-field-query/62526)

【问题讨论】:

    标签: search elasticsearch indexing n-gram elasticsearch-2.0


    【解决方案1】:

    我使用您的设置和映射来创建索引。你的查询似乎对我来说很好。我建议对正在返回的“意外”文档之一执行explain,看看为什么它被匹配并与其他结果一起返回。

    这是我所做的:

    在您的分析器上运行分析 api 以查看查询将如何拆分为标记。

    curl -XGET 'localhost:9200/my_index/_analyze' -d '
    {
      "analyzer" : "ngram_analyzer",
      "text" : "frodo"
    }'
    

    frodo 将使用您的分析器分成 3 个标记。

    {
      "tokens": [
        {
          "token": "fro",
          "start_offset": 0,
          "end_offset": 5,
          "type": "word",
          "position": 0
        },
        {
          "token": "rod",
          "start_offset": 0,
          "end_offset": 5,
          "type": "word",
          "position": 0
        },
        {
          "token": "odo",
          "start_offset": 0,
          "end_offset": 5,
          "type": "word",
          "position": 0
        }
      ]
    }
    

    我索引了 3 个文档进行测试(仅使用了 ngrams 字段)。以下是文档:

    {
      "took": 5,
      "timed_out": false,
      "_shards": {
        "total": 5,
        "successful": 5,
        "failed": 0
      },
      "hits": {
        "total": 3,
        "max_score": 1,
        "hits": [
          {
            "_index": "my_index",
            "_type": "my_type",
            "_id": "2",
            "_score": 1,
            "_source": {
              "ngram": "theodore"
            }
          },
          {
            "_index": "my_index",
            "_type": "my_type",
            "_id": "1",
            "_score": 1,
            "_source": {
              "ngram": "frodo"
            }
          },
          {
            "_index": "my_index",
            "_type": "my_type",
            "_id": "3",
            "_score": 1,
            "_source": {
              "ngram": "rudolpho"
            }
          }
        ]
      }
    }
    

    你提到的第一个查询,它匹配 frodo 和 theodore,但不匹配你提到的 rudolpho - 这是有道理的,因为 rudolpho 不会产生任何匹配 frodo 的 trigram 的 trigram

    frodo -> fro, rod, odo 
    
    rudolpho -> rud, udo, dol, olp, lph, pho
    

    使用您的第二个查询,我只返回 frodo (None of the other two) 。

    {
      "took": 5,
      "timed_out": false,
      "_shards": {
        "total": 5,
        "successful": 5,
        "failed": 0
      },
      "hits": {
        "total": 1,
        "max_score": 0.53148466,
        "hits": [
          {
            "_index": "my_index",
            "_type": "my_type",
            "_id": "1",
            "_score": 0.53148466,
            "_source": {
              "ngram": "frodo"
            }
          }
        ]
      }
    }
    

    然后我对其他两个文档(theodore 和 rudolpho)进行了解释 (localhost:9200/my_index/my_type/2/_explain),我看到了这个(我已经剪掉了回复)

    {
      "_index": "my_index",
      "_type": "my_type",
      "_id": "2",
      "matched": false,
      "explanation": {
        "value": 0,
        "description": "Failure to meet condition(s) of required/prohibited clause(s)",
        "details": [
          {
            "value": 0,
            "description": "no match on required clause ((ngram:fro ngram:rod ngram:odo)~2)",
            "details": [
    

    以上是预期的,因为来自 frodo 的三个令牌中至少有两个应该匹配。

    【讨论】:

    • 感谢您快速而彻底的回复!你对 rudolpho 是正确的,我打错了 rodolpho (是一个 fakerjs 生成的名字:).. 我要验证我只用我的索引的 ngram 查询得到相同的结果(我怀疑是这样),然后尝试添加查询名称字段并回复您。再次感谢!
    • 请参阅上面的要点和here 以了解匹配查询仍然返回每个匹配项和单个匹配三元组的解释。我仍在研究它,但老实说,_explain 结果有点令人困惑。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2015-04-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-11-26
    • 2018-02-04
    相关资源
    最近更新 更多