期望的行为是 ES 只返回至少包含
两个字符匹配。
我不熟悉 SmartCN analyzer,但如果您想匹配至少 2 个字符,那么根据您的用例,您可以使用 N-gram tokenizer,它在遇到以下情况时首先将文本分解为单词指定字符的列表,然后它发出指定长度的每个单词的 N-gram。
索引映射:
{
"settings": {
"analysis": {
"analyzer": {
"my_analyzer": {
"tokenizer": "my_tokenizer"
}
},
"tokenizer": {
"my_tokenizer": {
"type": "ngram",
"min_gram": 2, <-- note this
"max_gram": 20,
"token_chars": [
"letter",
"digit"
]
}
}
},
"max_ngram_diff": 50
},
"mappings": {
"properties": {
"title": {
"type": "text",
"analyzer": "my_analyzer",
"search_analyzer": "standard"
}
}
}
}
索引数据:
{
"title": "world"
}
分析 API
搜索查询将不匹配 "title": "w",因为生成的标记的最小长度为 2(因为 min_gram 在上面的索引映射中定义为 2)
生成的令牌是:
POST/_analyze
{
"tokens": [
{
"token": "wo",
"start_offset": 0,
"end_offset": 2,
"type": "word",
"position": 0
},
{
"token": "wor",
"start_offset": 0,
"end_offset": 3,
"type": "word",
"position": 1
},
{
"token": "worl",
"start_offset": 0,
"end_offset": 4,
"type": "word",
"position": 2
},
{
"token": "world",
"start_offset": 0,
"end_offset": 5,
"type": "word",
"position": 3
},
{
"token": "or",
"start_offset": 1,
"end_offset": 3,
"type": "word",
"position": 4
},
{
"token": "orl",
"start_offset": 1,
"end_offset": 4,
"type": "word",
"position": 5
},
{
"token": "orld",
"start_offset": 1,
"end_offset": 5,
"type": "word",
"position": 6
},
{
"token": "rl",
"start_offset": 2,
"end_offset": 4,
"type": "word",
"position": 7
},
{
"token": "rld",
"start_offset": 2,
"end_offset": 5,
"type": "word",
"position": 8
},
{
"token": "ld",
"start_offset": 3,
"end_offset": 5,
"type": "word",
"position": 9
}
]
}
**Search Query:**
{
"query": {
"match": {
"title": "wo"
}
}
}
搜索结果:
"hits": [
{
"_index": "stof_64003025",
"_type": "_doc",
"_id": "2",
"_score": 0.56802315,
"_source": {
"title": "world"
}
}
]