【问题标题】:elasticsearch mapping tokenizer keyword to avoid splitting tokens and enable use of wildcardelasticsearch 映射分词器关键字以避免拆分令牌并启用通配符
【发布时间】:2014-10-21 11:53:23
【问题描述】:

我尝试使用 angularjs 和 elasticsearch 在给定字段上创建一个自动完成功能,例如countryname。它可以包含简单的名称,如“法国”、“西班牙”或“组合名称”,如“塞拉利昂”。

在映射中,该字段为not_analyzed,以防止弹性标记“组合名称”

"COUNTRYNAME" : {"type" : "string", "store" : "yes","index": "not_analyzed" }

我需要查询elasticsearch:

  • 使用诸如“countryname:value”之类的内容过滤文档,其中 value 可以包含通配符
  • 并对过滤器返回的国家/地区名称进行聚合,(我进行聚合以仅获取不同的数据,计数对我来说没用,也许有更好的解决方案)

我不能在“not_analyzed”字段中使用通配符:

这是我的查询,但“值”变量中的通配符不起作用并且区分大小写:

通配符单独她的工作:

curl -XGET 'local_host:9200/botanic/specimens/_search?size=0' -d '{
  "fields": [
    "COUNTRYNAME"
  ],
  "query": {
    "query_string": {
      "query": "COUNTRYNAME:*"
    }
  },
  "aggs": {
    "general": {
      "terms": {
        "field": "COUNTRYNAME",
        "size": 0
      }
    }
  }
}'

但这不起作用(法郎*):

curl -XGET 'local_host:9200/botanic/specimens/_search?size=0' -d '{
  "fields": [
    "COUNTRYNAME"
  ],
  "query": {
    "query_string": {
      "query": "COUNTRYNAME:Franc*"
    }
  },
  "aggs": {
    "general": {
      "terms": {
        "field": "COUNTRYNAME",
        "size": 0
      }
    }
  }
}'

我也尝试使用bool must query,但不使用这个 not_analyzed 字段和通配符:

curl -XGET 'local_host:9200/botanic/specimens/_search?size=0' -d '{
  "fields": [
    "COUNTRYNAME"
  ],
  "query": {
    "bool": {
      "must": [
        {
          "match": {
            "COUNTRYNAME": "Franc*"
          }
        }
      ]
    }
  },
  "aggs": {
    "general": {
      "terms": {
        "field": "COUNTRYNAME",
        "size": 0
      }
    }
  }
}'

我错过了什么或做错了什么?我应该将字段analyzed 留在映射中并使用另一个不会将组合名称拆分为令牌的分析器吗??

【问题讨论】:

    标签: elasticsearch


    【解决方案1】:

    我找到了一个可行的解决方案:“关键字”标记器。 创建一个自定义分析器并将其用于我想要保留的字段的映射中,而不用空格分割:

        curl -XPUT 'localhost:9200/botanic/' -d '{
     "settings":{
         "index":{
            "analysis":{
               "analyzer":{
                  "keylower":{
                     "tokenizer":"keyword",
                     "filter":"lowercase"
                  }
               }
            }
         }
      },
      "mappings":{
            "specimens" : {
                "_all" : {"enabled" : true},
                "_index" : {"enabled" : true},
                "_id" : {"index": "not_analyzed", "store" : false},
                "properties" : {
                    "_id" : {"type" : "string", "store" : "no","index": "not_analyzed"  } ,
                ...
                    "LOCATIONID" : {"type" : "string",  "store" : "yes","index": "not_analyzed" } ,
                    "AVERAGEALTITUDEROUNDED" : {"type" : "string",  "store" : "yes","index": "analyzed" } ,
                    "CONTINENT" : {"type" : "string","analyzer":"keylower" } ,
                    "COUNTRYNAME" : {"type" : "string","analyzer":"keylower" } ,                
                    "COUNTRYCODE" : {"type" : "string", "store" : "yes","index": "analyzed" } ,
                    "COUNTY" : {"type" : "string","analyzer":"keylower" } ,
                    "LOCALITY" : {"type" : "string","analyzer":"keylower" }                 
                }
            }
        }
    }'
    

    所以我可以在字段 COUNTRYNAME 的查询中使用通配符,该字段未拆分:

    curl -XGET 'localhost:9200/botanic/specimens/_search?size=10' -d '{
    "fields"  : ["COUNTRYNAME"],     
    "query": {"query_string" : {
                        "query": "COUNTRYNAME:bol*"
    }},
    "aggs" : {
        "general" : {
            "terms" : {
                "field" : "COUNTRYNAME", "size":0
            }
        }
    }}'
    

    结果:

    {
        "took" : 14,
        "timed_out" : false,
        "_shards" : {
            "total" : 5,
            "successful" : 5,
            "failed" : 0
        },
        "hits" : {
            "total" : 45,
            "max_score" : 1.0,
            "hits" : [{
                    "_index" : "botanic",
                    "_type" : "specimens",
                    "_id" : "91E7B53B61DF4E76BF70C780315A5DFD",
                    "_score" : 1.0,
                    "fields" : {
                        "COUNTRYNAME" : ["Bolivia, Plurinational State of"]
                    }
                }, {
                    "_index" : "botanic",
                    "_type" : "specimens",
                    "_id" : "7D811B5D08FF4F17BA174A3D294B5986",
                    "_score" : 1.0,
                    "fields" : {
                        "COUNTRYNAME" : ["Bolivia, Plurinational State of"]
                    }
                } ...
            ]
        },
        "aggregations" : {
            "general" : {
                "buckets" : [{
                        "key" : "bolivia, plurinational state of",
                        "doc_count" : 45
                    }
                ]
            }
        }
    }
    

    【讨论】:

    • 感谢@Anainlb 的解决方案。这适用于我的类似案例。
    • 很高兴有帮助。随时分享任何改进;)
    • @AlainIb 这是我的错,因为您可以看到 ng-model 应该在 <input/> 标签内,但它在 <label> 标签内,甚至在您粘贴的答案中。请更新它。它现在工作..你为什么删除你的答案?我正要接受它
    • @Satyadev 我认为你发错了问题没有?
    • 太棒了!谢谢。我在 ES IRC,很难得到帮助。这篇文章为我节省了很多等待的时间,希望有人能帮助我!谢谢!!!!
    猜你喜欢
    • 2015-04-08
    • 1970-01-01
    • 2016-10-30
    • 2019-05-04
    • 2018-11-17
    • 2013-12-15
    • 2015-10-09
    • 2016-10-14
    • 2016-10-20
    相关资源
    最近更新 更多