【问题标题】:Searching names: increasing relevance for proximate matches in multivalued field搜索名称:增加多值字段中近似匹配的相关性
【发布时间】:2018-11-22 19:33:03
【问题描述】:

我正在尝试让名称字段在 Elasticsearch 中合理工作,但无法找到指导。请帮助我,互联网!

我的文档有多个作者,因此有一个多值名称字段。假设我要搜索 paul f tompkins,以及两个文档:{"authors": ["Paul Tompkins", "Dietrich Kohl"]} 和 {"authors": ["Paul Wang", "Darlene Tompkins"]}。

我的搜索将很容易地检索到这两个文档,但两者在authors 查询中的得分相同。我希望我在 authors 数组的同一项目中匹配多个术语以提高第一个文档的分数。

我该怎么做?我所知道的提高邻近度的两种技术是带状疱疹(我相信这会生成 paul_f 和 f_tompkins 带状疱疹,两者都不匹配)和带有 slop 的短语查询(由于 f 令牌而失败不存在)。

理想情况下,我想要一个类似于带有minimum_should_match 的短语 slop 查询:我给它四个单词,如果在同一个数组元素中至少存在两个单词,它匹配,并且在同一个数组元素中增加每个额外的匹配项分数。我不知道该怎么做。

(如果客户端逻辑试图从查询中去除f 对我来说是行不通的——这是一个简化的示例,但假设我也希望能够处理类似的查询paul francis tompkins 或 paul f tompkins there will be blood.)

【问题讨论】:

    标签: elasticsearch search full-text-search


    【解决方案1】:

    两个文档得分相同的原因是作者字段是文本值数组。如果我们改变存储作者的方式,我们可以得到想要的结果。为此,让作者成为nested 类型。所以我们有以下映射:

    "mappings": {
      "_doc": {
        "properties": {
          "authors": {
            "type": "nested",
            "properties": {
              "name": {
                "type": "text",
                "fields": {
                  "raw": {
                    "type": "keyword"
                  }
                }
              }
            }
          }
        }
      }
    }
    

    注意:子字段 raw 可以用于其他一些场景,没有关系到解决方案。

    现在让我们按如下方式索引文档:

    文件 1:

    {
      "authors": [
        {
          "name": "Paul Tompkins"
        },
        {
          "name": "Dietrich Kohl"
        }
      ]
    }
    

    文件 2:

    {
      "authors": [
        {
          "name": "Paul Wang"
        },
        {
          "name": "Darlene Tompkins"
        }
      ]
    }
    

    让我们按如下方式查询它们:

    {
      "explain": true,
      "query": {
        "nested": {
          "path": "authors",
          "query": {
            "query_string": {
              "query": "paul l tompkins",
              "fields": [
                "authors.name"
              ]
            }
          }
        }
      }
    }
    

    结果:

      "hits": {
        "total": 2,
        "max_score": 1.3862944,
        "hits": [
          {
            "_index": "test",
            "_type": "_doc",
            "_id": "1",
            "_score": 1.3862944,
            "_source": {
              "authors": [
                {
                  "name": "Paul Tompkins"
                },
                {
                  "name": "Dietrich Kohl"
                }
              ]
            }
          },
          {
            "_index": "test",
            "_type": "_doc",
            "_id": "2",
            "_score": 0.6931472,
            "_source": {
              "authors": [
                {
                  "name": "Paul Wang"
                },
                {
                  "name": "Darlene Tompkins"
                }
              ]
            }
          }
        ]
      }
    

    注意:在查询中,我也使用了 explain:true。这给出了分数计算的解释。(我没有包括上面的解释输出,因为它很长。不过你可以试试)。

    当我们查看评分机制时,我们可以看到查询嵌套字段和查询数组时的区别。从广义上讲,由于嵌套字段存储为单独的文档,因此 Doc 1 得分较高,因为子文档 1 即:

    {
      "name": "Paul Tompkins"
    }
    

    将获得更高的分数,因为 paul 和 tompkins 这两个词都在同一个子文档中。

    在数组的情况下,所有名称都属于同一字段,而不是单独的子文档,因此存在差异。

    这样我们就可以达到预期的效果了。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-12-23
      • 2011-01-07
      • 2015-06-11
      • 1970-01-01
      相关资源
      最近更新 更多