【发布时间】:2021-11-10 09:43:52
【问题描述】:
我发现 SQL Server 全文搜索的一个非常奇怪的行为是索引 SUR、SCR 和可能的其他一些首字母缩略词,以及它后面的数字 - 作为“精确匹配”。
SELECT * FROM sys.dm_fts_parser ('"SUR 12345"', 1033, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| s u r 1 2 3 4 5 | 1 | 0 | 1 | Exact Match | sur 12345 | 0 | SUR 12345 |
| n n 1 2 3 4 5 s u r | 1 | 0 | 1 | Exact Match | nn12345sur | 0 | SUR 12345 |
SELECT * FROM sys.dm_fts_parser ('"SCR 12345"', 1033, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| s c r 1 2 3 4 5 | 1 | 0 | 1 | Exact Match | scr 12345 | 0 | SCR 12345 |
| n n 1 2 3 4 5 s c r | 1 | 0 | 1 | Exact Match | nn12345scr | 0 | SCR 12345 |
其他首字母缩写词或文本,包括小写 sur,不受影响:
SELECT * FROM sys.dm_fts_parser ('"sur 12345"', 1033, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| s u r | 1 | 0 | 1 | Exact Match | sur | 0 | sur 12345 |
| 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | 12345 | 0 | sur 12345 |
| n n 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | nn12345 | 0 | sur 12345 |
SELECT * FROM sys.dm_fts_parser ('"ABC 12345"', 1033, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| a b c | 1 | 0 | 1 | Exact Match | abc | 0 | ABC 12345 |
| 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | 12345 | 0 | ABC 12345 |
| n n 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | nn12345 | 0 | ABC 12345 |
SELECT * FROM sys.dm_fts_parser ('"XYZ 76"', 1033, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| x y z | 1 | 0 | 1 | Exact Match | xyz | 0 | XYZ 76 |
| 7 6 | 1 | 0 | 2 | Exact Match | 76 | 0 | XYZ 76 |
| n n 7 6 | 1 | 0 | 2 | Exact Match | nn76 | 0 | XYZ 76 |
这种行为似乎出乎意料,很可能是错误的,但我也可能遗漏了一些与断字有关的明显内容(尝试 1033 和 2057 - 效果相同)。我在 SQL Server 2019 Linux 15.0.4053.23 和 2017 CU20 和 CU25 上复制了它,我可以立即访问。
有没有人有类似的问题和解决方案,以便 SUR、SCR 和任何其他可能损坏的首字母缩写词将独立于以下数字进行索引?
编辑:
将语言更改为 0(中性)会导致奇怪的行为 - 它不能解决使用 SUR 首字母缩写词时的问题,但会修复 SCR 首字母缩写词!
SELECT * FROM sys.dm_fts_parser ('"SUR 12345"', 0, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| s u r 1 2 3 4 5 | 1 | 0 | 1 | Exact Match | sur 12345 | 0 | SUR 12345 |
| n n 1 2 3 4 5 s u r | 1 | 0 | 1 | Exact Match | nn12345sur | 0 | SUR 12345 |
SELECT * FROM sys.dm_fts_parser ('"SCR 12345"', 0, 0, 0)
| keyword | group_id | phrase_id | occurrence | special_term | display_term | expansion_type | source_term |
|---|---|---|---|---|---|---|---|
| s c r | 1 | 0 | 1 | Exact Match | scr | 0 | SCR 12345 |
| 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | 12345 | 0 | SCR 12345 |
| n n 1 2 3 4 5 | 1 | 0 | 2 | Exact Match | nn12345 | 0 | SCR 12345 |
我决定悬赏这个问题,因为理想情况下我需要通过重新配置数据库索引来解决找不到搜索词的问题。
为了帮助重现下面的问题,是一个创建数据库的脚本(带有注释掉的 DROP 脚本以帮助重置状态)
/*
DROP FULLTEXT INDEX ON EnglishTexts
DROP FULLTEXT INDEX ON NeutralTexts
DROP FULLTEXT CATALOG TestSearchCatalog
USE master
DROP DATABASE TestSearch
*/
CREATE DATABASE TestSearch
GO
USE [TestSearch]
GO
CREATE FULLTEXT CATALOG TestSearchCatalog WITH ACCENT_SENSITIVITY = OFF
GO
CREATE TABLE EnglishTexts (Id INT IDENTITY(1,1) NOT NULL, Text NVARCHAR(MAX), CONSTRAINT PK_EnglishTexts PRIMARY KEY CLUSTERED (Id))
CREATE FULLTEXT INDEX ON EnglishTexts (Text LANGUAGE 'English') KEY INDEX PK_EnglishTexts ON ([TestSearchCatalog]) WITH (CHANGE_TRACKING = AUTO, STOPLIST = OFF)
INSERT INTO EnglishTexts(Text) VALUES ('PRFX 12233')
INSERT INTO EnglishTexts(Text) VALUES ('SUR 12233')
INSERT INTO EnglishTexts(Text) VALUES ('SCR 12233')
CREATE TABLE NeutralTexts (Id INT IDENTITY(1,1) NOT NULL, Text NVARCHAR(MAX), CONSTRAINT PK_NeutralTexts PRIMARY KEY CLUSTERED (Id))
CREATE FULLTEXT INDEX ON NeutralTexts (Text LANGUAGE 'Neutral') KEY INDEX PK_NeutralTexts ON ([TestSearchCatalog]) WITH (CHANGE_TRACKING = AUTO, STOPLIST = OFF)
INSERT INTO NeutralTexts(Text) VALUES ('PRFX 12233')
INSERT INTO NeutralTexts(Text) VALUES ('SUR 12233')
INSERT INTO NeutralTexts(Text) VALUES ('SCR 12233')
-- following query returns 1 row but should 3 - a possible bug in english word breaker
SELECT * FROM EnglishTexts WHERE CONTAINS(Text, '"12233"')
-- following query returns 2 rows but should 3 - neutral language word breaker is also treating SUR acronym specially - another bug?
SELECT * FROM NeutralTexts WHERE CONTAINS(Text, '"12233"')
-- following query returns 1 row but should 3 - forcing neutral language on a query on english index should apply neutral language (i might misunderstand if this is even possible without a neutral index)
SELECT * FROM EnglishTexts WHERE CONTAINS(Text, '"12233"', LANGUAGE 0)
-- following query returns 2 rows but should 3 - using neutral language on neutral language indexed table should not make a difference
SELECT * FROM NeutralTexts WHERE CONTAINS(Text, '"12233"', LANGUAGE 0)
-- for reference - English word breaker does not split SCR with 12233 and SUR with 12233, causing above problems
SELECT * FROM sys.dm_fts_parser ('"SCR 12233 SUR 12233"', 1033, 0, 0)
-- for reference - Neutral word breaker correctly splits SCR and 12233 but not SUR with 12233
SELECT * FROM sys.dm_fts_parser ('"SCR 12233 SUR 12233"', 0, 0, 0)
【问题讨论】:
-
尝试发音为
SCR。然后注意c在pronouncing中的发音。SCR听起来与SR相同,而SUR和SIR听起来相同。 -
全文搜索基本上是通过搜索单词和文本来创建的,您的示例对于全文来说是不合理的。您应该尝试创建一个新的停止列表(空),在该停止列表中的停止词中添加一个空格。使用此非索引字表重建您的全文索引。与丹麦语的答案相同,但只需创建一个新的空停止列表并在其中添加一个空格
-
另外,如果您解释您正在尝试的搜索类型,我可以提供帮助,它看起来像条形码编号。您是否尝试在带有条形码的产品上查找内容?
-
示例非常合理 - 文本包含首字母缩写词、空格和数字,例如“图书 ISBN 1234567 很好”。如果首字母缩写词是 SCR 或 SUR 而不是 ISBN,则搜索不会返回 1234567 的结果。
-
@MatBailie 关于发音,听起来(双关语)是一个很好的引导,尽管我认为中性语言不应该受到 SUR 的英语发音的影响
标签: sql sql-server parsing indexing full-text-search