假设这些样本数据 - 请参阅下面的 CTAS 语句来创建它们。
FIRSTNAME LASTNAME
--------- --------
D'Arch O'Neil
DArch ONeil
D Arch O Neil
D.Arch O.Neil
要保留撇号,您必须将其定义为printjoin 字符。否则将在分词过程中被删除(与最后一行中的点比较)。
BEGIN
ctxsys.ctx_ddl.create_preference('lex', 'BASIC_LEXER');
ctxsys.ctx_ddl.set_attribute('lex', 'printjoins', '''');
END;
/
---
BEGIN
ctxsys.ctx_ddl.create_preference('pref', 'MULTI_COLUMN_DATASTORE');
ctx_ddl.set_attribute('pref', 'columns', 'firstname, lastname');
END;
/
create index idx_test on test(firstname)
indextype is CTXSYS.CONTEXT
parameters ('datastore pref LEXER lex')
;
创建索引后查看结果tokes的最佳方法是查询包含所有tokes的$I表:
select TOKEN_TEXT from DR$IDX_TEST$I;
TOKEN_TEXT
----------------------------------------------------------------
ARCH
D'ARCH
DARCH
FIRSTNAME
LASTNAME
NEIL
O
O'NEIL
ONEIL
正如预期的那样,点消失了,但撇号仍然存在。
现在你可以查询并且得到预期的结果:
select *
from test
where CONTAINS (firstname, '%D''Ar%') >0;
FIRSTNAME LASTNAME
--------- --------
D'Arch O'Neil
当然,当您使用MULTI_COLUMN_DATASTORE 时,您还可以查询lastname 列中的数据。
select *
from test
where CONTAINS (firstname, '%O''Ne%') >0;
FIRSTNAME LASTNAME
--------- --------
D'Arch O'Neil
样本数据
create table test as
select 'D''Arch' firstname, 'O''Neil' Lastname from dual union all
select 'DArch' firstname, 'ONeil' Lastname from dual union all
select 'D Arch' firstname, 'O Neil' Lastname from dual union all
select 'D.Arch' firstname, 'O.Neil' Lastname from dual;