【发布时间】:2013-03-14 11:13:34
【问题描述】:
我正在研究 Lucene Java 算法。 我们在 MySQL 数据库中有 100K 停止名称。 站点名称就像
NEW YORK PENN STATION,
NEWARK PENN STATION,
NEWARK BROAD ST,
NEW PROVIDENCE
etc
当用户输入像 NEW YORK 这样的搜索输入时,我们会在结果中得到 NEW YORK PENN STATION,但当用户输入准确的 NEW YORK PENN STATION 在搜索输入中返回零个结果。
我的代码是 -
public ArrayList<String> getSimilarString(ArrayList<String> source, String querystr)
{
ArrayList<String> arResult = new ArrayList<String>();
try
{
// 0. Specify the analyzer for tokenizing text.
// The same analyzer should be used for indexing and searching
StandardAnalyzer analyzer = new StandardAnalyzer(Version.LUCENE_40);
// 1. create the index
Directory index = new RAMDirectory();
IndexWriterConfig config = new IndexWriterConfig(Version.LUCENE_40, analyzer);
IndexWriter w = new IndexWriter(index, config);
for(int i = 0; i < source.size(); i++)
{
addDoc(w, source.get(i), "1933988" + (i + 1) + "z");
}
w.close();
// 2. query
// the "title" arg specifies the default field to use
// when no field is explicitly specified in the query.
Query q = new QueryParser(Version.LUCENE_40, "title", analyzer).parse(querystr + "*");
// 3. search
int hitsPerPage = 20;
IndexReader reader = DirectoryReader.open(index);
IndexSearcher searcher = new IndexSearcher(reader);
TopScoreDocCollector collector = TopScoreDocCollector.create(hitsPerPage, true);
searcher.search(q, collector);
ScoreDoc[] hits = collector.topDocs().scoreDocs;
// 4. Get results
for(int i = 0; i < hits.length; ++i)
{
int docId = hits[i].doc;
Document d = searcher.doc(docId);
arResult.add(d.get("title"));
}
// reader can only be closed when there
// is no need to access the documents any more.
reader.close();
}
catch(Exception e)
{
System.out.println("Exception (LuceneAlgo.getSimilarString()) : " + e);
}
return arResult;
}
private static void addDoc(IndexWriter w, String title, String isbn) throws IOException
{
Document doc = new Document();
doc.add(new TextField("title", title, Field.Store.YES));
// use a string field for isbn because we don't want it tokenized
doc.add(new StringField("isbn", isbn, Field.Store.YES));
w.addDocument(doc);
}
在此代码中,source 是停止名称列表,query 是用户给定的搜索输入。
Lucene 算法是否适用于大字符串?
为什么 Lucene 算法不适用于 Exact String?
【问题讨论】:
-
不确定答案,但您是否尝试过使用 Luke 查看您的索引? getopt.org/luke 可能是索引中的某些内容被篡改了。
-
谢谢@Taylor,我会试试这个
-
“不工作”是什么意思?您的查询和返回结果是什么?
-
您正在指定与停止名称相关的域,但代码中包含像“title”和“isbn”这样的字段名称,我真的很好奇..
标签: java lucene string-matching