【发布时间】:2012-09-16 19:04:51
【问题描述】:
我需要按姓名搜索人员。这里的人名可以是英文、韩文或中文。为此,我使用Like 条件在Name 的基础上进行搜索,如下所示:
select * from [MyTable] where Name like N'%t%'
上面的声明是给所有包含t的用户。但这不适用于韩语或中文。就像我用韩文字母 ㅈ 搜索一样,它应该给出包含这个字母的所有名称,例如 **정수연, 재훈아이팟, 정원혁 테스트 7**。我尝试了以下方法,但结果为零
select * from [MyTable] where Name like N'%ㅈ%' - No Results
select PATINDEX(N'%ㅈ%',N'정수연(Mohan)') - giving value as ZERO
select Charindex(N'ㅈ',N'정수연') - giving value as ZERO
有没有办法在SQL server中查找其他语言的字母表?
我知道如何使用编码技术在 C# 单词中找到其他语言中存在的字母,但在 SQL Server 中却不知道。请在这方面帮助我。
提前致谢。
编辑 C# 代码
public static string DecomposeSyllabels(string unicodeString) {
try {
//Consonant consonant only used
string[] JLT = { "ㄱ", "ㄲ", "ㄴ", "ㄷ", "ㄸ", "ㄹ", "ㅁ", "ㅂ", "ㅃ", "ㅅ", "ㅆ", "ㅇ", "ㅈ", "ㅉ", "ㅊ", "ㅋ", "ㅌ", "ㅍ", "ㅎ" };
// Only used a collection of neutral
string[] JVT = { "ㅏ", "ㅐ", "ㅑ", "ㅒ", "ㅓ", "ㅔ", "ㅕ", "ㅖ", "ㅗ", "ㅘ", "ㅙ", "ㅚ", "ㅛ", "ㅜ", "ㅝ", "ㅞ", "ㅟ", "ㅠ", "ㅡ", "ㅢ", "ㅣ" };
// Initial and coda consonants used in
string[] JTT = { "", "ㄱ", "ㄲ", "ㄳ", "ㄴ", "ㄵ", "ㄶ", "ㄷ", "ㄹ", "ㄺ", "ㄻ", "ㄼ", "ㄽ", "ㄾ", "ㄿ", "ㅀ", "ㅁ", "ㅂ", "ㅄ", "ㅅ", "ㅆ", "ㅇ", "ㅈ", "ㅊ", "ㅋ", "ㅌ", "ㅍ", "ㅎ" };
double SBase = 0xAC00;
long SCount = 11172;
int TCount = 28;
int NCount = 588;
string syllables = string.Empty;
foreach (char c in unicodeString) {
double SIndex = (int)c - SBase;
if (0 > SIndex || SIndex >= SCount) {
syllables = syllables + c;
continue;
}
int LIndex = (int)Math.Floor(SIndex / NCount);
int VIndex = (int)(Math.Floor((SIndex % NCount) / TCount));
int TIndex = (int)(SIndex % TCount);
syllables = syllables + (JLT[LIndex] + JVT[VIndex] + JTT[TIndex]);
}
return syllables;
}
catch {
return unicodeString;
}
}
【问题讨论】:
-
问题是 PATINDEX 和 CHARINDEX 搜索字符。而
ㅈ可以被认为是정的一部分;在 Unicode 级别U+3148不是U+C815的“部分” - 它们是单独的字符。 -
这个问题与Unicode规范化有关。在除 Macintosh 之外的大多数平台上,韩语由音节块编码。 Mac 按字母编码,渲染器将它们排列成音节块。使用编程语言,您将能够找到一些代码来进行
NFC和NFD规范化,但是使用SQL,我不知道是否存在这样的功能。可能值得添加特殊列。
标签: sql-server unicode sql-server-2008-r2 cjk