【问题标题】:How to prevent Elasticsearch from matching on only one non-English character with multi_match如何使用 multi_match 防止 Elasticsearch 仅匹配一个非英文字符
【发布时间】:2020-09-22 04:18:15
【问题描述】:

我正在使用elasiticsearch-dsl 和 SmartCN 分析器:

from elasticsearch_dsl import analyzer
analyzer_cn = analyzer(
    'smartcn',
    tokenizer=tokenizer('smartcn_tokenizer'),
    filter=['lowercase']
)

我正在使用 multi_match 来匹配多个术语:

from elasticsearch_dsl import Q
q_new = Q("multi_match", query="SOME_QUERY", fields=["FEIDL_NAME"])

期望的行为是 ES 只返回至少匹配两个字符的文档。我查看了文档,但找不到阻止 Elasticsearch 匹配单个字符的方法。

高度赞赏任何方向/建议。 谢谢。

【问题讨论】:

  • 你有没有机会浏览我的答案,期待得到你的反馈:)

标签: python-3.x elasticsearch elasticsearch-dsl


【解决方案1】:

期望的行为是 ES 只返回至少包含 两个字符匹配。

我不熟悉 SmartCN analyzer,但如果您想匹配至少 2 个字符,那么根据您的用例,您可以使用 N-gram tokenizer,它在遇到以下情况时首先将文本分解为单词指定字符的列表,然后它发出指定长度的每个单词的 N-gram。

索引映射:

{
    "settings": {
        "analysis": {
            "analyzer": {
                "my_analyzer": {
                    "tokenizer": "my_tokenizer"
                }
            },
            "tokenizer": {
                "my_tokenizer": {
                    "type": "ngram",
                    "min_gram": 2,    <-- note this
                    "max_gram": 20,
                    "token_chars": [
                        "letter",
                        "digit"
                    ]
                }
            }
        },
        "max_ngram_diff": 50
    },
    "mappings": {
        "properties": {
            "title": {
                "type": "text",
                "analyzer": "my_analyzer",
                "search_analyzer": "standard"
            }
        }
    }
}

索引数据:

{
    "title": "world"
}

分析 API

搜索查询将不匹配 "title": "w",因为生成的标记的最小长度为 2(因为 min_gram 在上面的索引映射中定义为 2)

生成的令牌是:

POST/_analyze

{
  "tokens": [
    {
      "token": "wo",
      "start_offset": 0,
      "end_offset": 2,
      "type": "word",
      "position": 0
    },
    {
      "token": "wor",
      "start_offset": 0,
      "end_offset": 3,
      "type": "word",
      "position": 1
    },
    {
      "token": "worl",
      "start_offset": 0,
      "end_offset": 4,
      "type": "word",
      "position": 2
    },
    {
      "token": "world",
      "start_offset": 0,
      "end_offset": 5,
      "type": "word",
      "position": 3
    },
    {
      "token": "or",
      "start_offset": 1,
      "end_offset": 3,
      "type": "word",
      "position": 4
    },
    {
      "token": "orl",
      "start_offset": 1,
      "end_offset": 4,
      "type": "word",
      "position": 5
    },
    {
      "token": "orld",
      "start_offset": 1,
      "end_offset": 5,
      "type": "word",
      "position": 6
    },
    {
      "token": "rl",
      "start_offset": 2,
      "end_offset": 4,
      "type": "word",
      "position": 7
    },
    {
      "token": "rld",
      "start_offset": 2,
      "end_offset": 5,
      "type": "word",
      "position": 8
    },
    {
      "token": "ld",
      "start_offset": 3,
      "end_offset": 5,
      "type": "word",
      "position": 9
    }
  ]
}

**Search Query:**

    {
        "query": {
            "match": {
                "title": "wo"
            }
        }
    }

搜索结果:

"hits": [
      {
        "_index": "stof_64003025",
        "_type": "_doc",
        "_id": "2",
        "_score": 0.56802315,
        "_source": {
          "title": "world"
        }
      }
    ]

【讨论】:

  • 您好 Bhavya,感谢您的回复。这种方法不起作用,因为我们需要使用 SmartCn 来标记化非英文文本。是否可以在查询级别而不是在标记化级别限制全文搜索的长度?
猜你喜欢
  • 2013-11-15
  • 2013-05-31
  • 1970-01-01
  • 1970-01-01
  • 2017-03-02
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多