【问题标题】:ElasticSearch - return the complete value of a facet for a queryElasticSearch - 为查询返回一个方面的完整值
【发布时间】:2014-01-27 15:34:09
【问题描述】:

我最近开始使用 ElasticSearch。我尝试完成一些用例。我对其中一个有疑问。

我已经用他们的全名(例如“Jean-Paul Gautier”、“Jean De La Fontaine”)索引了一些用户。

我尝试获取所有响应某个查询的全名。

例如,我想要以“J”开头的 100 个最常见的全名

{
  "query": {
    "query_string" : { "query": "full_name:J*" } }
  },
  "facets":{
    "name":{
      "terms":{
        "field": "full_name",
        "size":100
      }
    }
  }
}

我得到的结果是全名的所有单词:“Jean”、“Paul”、“Gautier”、“De”、“La”、“Fontaine”。

如何获得“Jean-Paul Gautier”和“Jean De La Fontaine”(所有全名值都由“J”求出)? “post_filter”选项没有这样做,它只限制了上面的子集。

  • 我必须配置“如何工作”这个全名方面
  • 我必须在当前查询中添加一些选项
  • 我必须做一些“映射”(目前非常模糊)

谢谢

【问题讨论】:

    标签: lucene elasticsearch


    【解决方案1】:

    您只需在该字段上设置"index": "not_analyzed",您就可以在您的构面中取回完整的、未修改的字段值。

    通常,最好有一个不分析的字段版本(用于分面)和另一个版本(用于搜索)。 "multi_field" 字段类型对此很有用。

    所以在这种情况下,我可以定义如下映射:

    curl -XPUT "http://localhost:9200/test_index/" -d'
    {
       "mappings": {
          "people": {
             "properties": {
                "full_name": {
                   "type": "multi_field",
                   "fields": {
                      "untouched": {
                         "type": "string",
                         "index": "not_analyzed"
                      },
                      "full_name": {
                         "type": "string"
                      }
                   }
                }
             }
          }
       }
    }'
    

    这里我们有两个子字段。与父级同名的将是默认值,因此如果您搜索"full_name" 字段,Elasticsearch 将实际使用"full_name.full_name"。 "full_name.untouched" 会给你你想要的方面结果。

    所以接下来我添加两个文档:

    curl -XPUT "http://localhost:9200/test_index/people/1" -d'
    {
       "full_name": "Jean-Paul Gautier"
    }'
    
    curl -XPUT "http://localhost:9200/test_index/people/2" -d'
    {
       "full_name": "Jean De La Fontaine"
    }'
    

    然后我可以对每个字段进行分面以查看返回的内容:

    curl -XPOST "http://localhost:9200/test_index/_search" -d'
    {
       "size": 0,
       "facets": {
          "name_terms": {
             "terms": {
                "field": "full_name"
             }
          },
          "name_untouched": {
             "terms": {
                "field": "full_name.untouched",
                "size": 100
             }
          }
       }
    }'
    

    我得到以下信息:

    {
       "took": 1,
       "timed_out": false,
       "_shards": {
          "total": 1,
          "successful": 1,
          "failed": 0
       },
       "hits": {
          "total": 2,
          "max_score": 0,
          "hits": []
       },
       "facets": {
          "name_terms": {
             "_type": "terms",
             "missing": 0,
             "total": 7,
             "other": 0,
             "terms": [
                {
                   "term": "jean",
                   "count": 2
                },
                {
                   "term": "paul",
                   "count": 1
                },
                {
                   "term": "la",
                   "count": 1
                },
                {
                   "term": "gautier",
                   "count": 1
                },
                {
                   "term": "fontaine",
                   "count": 1
                },
                {
                   "term": "de",
                   "count": 1
                }
             ]
          },
          "name_untouched": {
             "_type": "terms",
             "missing": 0,
             "total": 2,
             "other": 0,
             "terms": [
                {
                   "term": "Jean-Paul Gautier",
                   "count": 1
                },
                {
                   "term": "Jean De La Fontaine",
                   "count": 1
                }
             ]
          }
       }
    }
    

    如您所见,已分析字段返回单个单词、小写标记(当您未指定分析器时,使用standard analyzer),未分析子字段返回未修改的原始标记文本。

    这是一个您可以使用的可运行示例: http://sense.qbox.io/gist/7abc063e2611846011dd874648fd1b77450b19a5

    【讨论】:

    • 在更新当前映射遇到一些困难后,我成功了!感谢您的宝贵帮助。
    【解决方案2】:

    尝试更改“full_name”的映射:

    "properties": {
      "full_name": {
         "type": "string",
         "index": "not_analyzed"
      }
      ...
    }
    

    not_analyzed 表示将保持原样,包括大写字母、空格、破折号等,以便“Jean De La Fontaine”保持可查找性并且不会被标记为“Jean”“De”“La”“Fontaine”

    你可以experiment with different analyzers using the api

    注意标准对多部分名称的作用:

    GET /_analyze?analyzer=standard
    {'Jean Claude Van Dame'}
    
    
    {
       "tokens": [
          {
             "token": "jean",
             "start_offset": 2,
             "end_offset": 6,
             "type": "<ALPHANUM>",
             "position": 1
          },
          {
             "token": "claude",
             "start_offset": 7,
             "end_offset": 13,
             "type": "<ALPHANUM>",
             "position": 2
          },
          {
             "token": "van",
             "start_offset": 14,
             "end_offset": 17,
             "type": "<ALPHANUM>",
             "position": 3
          },
          {
             "token": "dame",
             "start_offset": 18,
             "end_offset": 22,
             "type": "<ALPHANUM>",
             "position": 4
          }
       ]
    }
    

    【讨论】:

    • 感谢“分析器”的链接,这很有用!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2020-03-25
    • 1970-01-01
    • 1970-01-01
    • 2012-04-30
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多