您只需在该字段上设置"index": "not_analyzed",您就可以在您的构面中取回完整的、未修改的字段值。
通常,最好有一个不分析的字段版本(用于分面)和另一个版本(用于搜索)。 "multi_field" 字段类型对此很有用。
所以在这种情况下,我可以定义如下映射:
curl -XPUT "http://localhost:9200/test_index/" -d'
{
"mappings": {
"people": {
"properties": {
"full_name": {
"type": "multi_field",
"fields": {
"untouched": {
"type": "string",
"index": "not_analyzed"
},
"full_name": {
"type": "string"
}
}
}
}
}
}
}'
这里我们有两个子字段。与父级同名的将是默认值,因此如果您搜索"full_name" 字段,Elasticsearch 将实际使用"full_name.full_name"。 "full_name.untouched" 会给你你想要的方面结果。
所以接下来我添加两个文档:
curl -XPUT "http://localhost:9200/test_index/people/1" -d'
{
"full_name": "Jean-Paul Gautier"
}'
curl -XPUT "http://localhost:9200/test_index/people/2" -d'
{
"full_name": "Jean De La Fontaine"
}'
然后我可以对每个字段进行分面以查看返回的内容:
curl -XPOST "http://localhost:9200/test_index/_search" -d'
{
"size": 0,
"facets": {
"name_terms": {
"terms": {
"field": "full_name"
}
},
"name_untouched": {
"terms": {
"field": "full_name.untouched",
"size": 100
}
}
}
}'
我得到以下信息:
{
"took": 1,
"timed_out": false,
"_shards": {
"total": 1,
"successful": 1,
"failed": 0
},
"hits": {
"total": 2,
"max_score": 0,
"hits": []
},
"facets": {
"name_terms": {
"_type": "terms",
"missing": 0,
"total": 7,
"other": 0,
"terms": [
{
"term": "jean",
"count": 2
},
{
"term": "paul",
"count": 1
},
{
"term": "la",
"count": 1
},
{
"term": "gautier",
"count": 1
},
{
"term": "fontaine",
"count": 1
},
{
"term": "de",
"count": 1
}
]
},
"name_untouched": {
"_type": "terms",
"missing": 0,
"total": 2,
"other": 0,
"terms": [
{
"term": "Jean-Paul Gautier",
"count": 1
},
{
"term": "Jean De La Fontaine",
"count": 1
}
]
}
}
}
如您所见,已分析字段返回单个单词、小写标记(当您未指定分析器时,使用standard analyzer),未分析子字段返回未修改的原始标记文本。
这是一个您可以使用的可运行示例:
http://sense.qbox.io/gist/7abc063e2611846011dd874648fd1b77450b19a5