【发布时间】:2017-02-16 21:33:26
【问题描述】:
背景:我已通过索引标记化名称(name 字段)以及三元组分析名称(ngram 字段)对名称字段进行了部分搜索。
我已经提升了 name 字段以使精确的标记匹配冒泡到结果的顶部。
问题:我正在尝试实现一个查询,将 nGram 匹配限制为仅匹配查询字符串的某个阈值(例如 80%)的查询。我知道minimum_should_match 似乎是我正在寻找的东西,但我的问题是形成查询以实际产生这些结果。
我的确切标记匹配被提升到顶部,但我仍然得到 每个在 ngram 字段中具有单个匹配三元组的文档。
GIST: Index settings and mapping
索引设置
{
"my_index": {
"settings": {
"index": {
"number_of_shards": "5",
"max_result_window": "30000",
"creation_date": "1475853851937",
"analysis": {
"filter": {
"ngram_filter": {
"type": "ngram",
"min_gram": "3",
"max_gram": "3"
}
},
"analyzer": {
"ngram_analyzer": {
"filter": [
"lowercase",
"ngram_filter"
],
"type": "custom",
"tokenizer": "standard"
}
}
},
"number_of_replicas": "1",
"uuid": "AuCjcP5sSb-m59bYrprFcw",
"version": {
"created": "2030599"
}
}
}
}
}
索引映射
{
"my_index": {
"mappings": {
"my_type": {
"properties": {
"acw": {
"type": "integer"
},
"pcg": {
"type": "integer"
},
"date": {
"type": "date",
"format": "strict_date_optional_time||epoch_millis"
},
"dob": {
"type": "date",
"format": "strict_date_optional_time||epoch_millis"
},
"id": {
"type": "string"
},
"name": {
"type": "string",
"boost": 10
},
"ngram": {
"type": "string",
"analyzer": "ngram_analyzer"
},
"bdk": {
"type": "integer"
},
"mmw": {
"type": "integer"
},
"mpi": {
"type": "integer"
},
"sex": {
"type": "string",
"index": "not_analyzed"
}
}
}
}
}
}
解决方案尝试
[GIST:查询尝试] 由于 2 个链接限制而取消链接:(
(https://gist.github.com/jordancardwell/2e690013666e7e1da6ef1acee314b4e6)
我尝试了一个多重匹配查询,它给了我正确的搜索结果,但我没有运气省略只匹配单个三元组的名称的结果(比如“odo”三元组里面“ odophilus")
//this matches 'frodo' and sends results to the top, since `name` field is boosted
// but also matches 'theodore' and 'rodolpho'
{
"size":100,
"from":0,
"query":{
"multi_match":{
"query":"frodo",
"fields":[
"name",
"ngram"
],
"type":"best_fields"
}
}
}
.
//I then tried to throw in the `minimum_must_match` option
// hoping it would filter out large strings that only had one matching trigram for instance
{
"size":100,
"from":0,
"query":{
"multi_match":{
"query":"frodo",
"fields":[
"name",
"ngram"
],
"type":"best_fields",
"minimum_should_match": "90%",
}
}
}
我已经尝试过有意义地手动生成匹配查询,从而允许我仅将 minimum_must_match 应用于 ngram 字段,但似乎无法正确使用语法。
// I then tried to contruct a custom query to just return the `minimum_should_match`d results on the ngram field
// I started with a query produced by using bodybuilder to `and` and `or` my other search criteria together
{
"query": {
"bool": {
"filter": {
"bool": {
"must": [
//each separate field's criteria `must`/`and`ed together
{
"query": {
"bool": {
"filter": {
"bool": {
"should": [
//each critereon for a specific field `should`/`or`ed together
{
//my attempt at getting `ngram` field results..
// should theoretically only return when field
// contains nothing but matching ngrams
// (i.e. exact matches and other fluke matches)
"query": {
"match": {
"ngram": {
"query": "frodo",
"minimum_should_match": "100%"
}
}
}
}
//... other critereon to be `should`/`or`ed together
]
}
}
}
}
}
//... other criteria to be `must`/`and`ed together
]
}
}
}
}
}
谁能看出我做错了什么?
看起来这应该很容易完成,但我肯定遗漏了一些明显的东西。
更新
我使用 _explain=true(使用 sense UI)运行了一个查询,试图了解我的结果。
我在"frod" 的"frod" 和minimum_should_match = 100% 的ngram 字段上查询了match,但我仍然得到与至少一个ngram 匹配的每条记录。
(例如rodolpho,即使它不包含fro)
注意:交叉发布自 [discuss.elastic.co] 稍后会做一个链接,不能发布超过2个:/
(https://discuss.elastic.co/t/ngram-partial-match-limiting-ngram-results-in-multiple-field-query/62526)
【问题讨论】:
标签: search elasticsearch indexing n-gram elasticsearch-2.0