【问题标题】:Why Lucene algorithm not working for Exact String in Java?为什么 Lucene 算法不适用于 Java 中的精确字符串?
【发布时间】:2013-03-14 11:13:34
【问题描述】:

我正在研究 Lucene Java 算法。 我们在 MySQL 数据库中有 100K 停止名称。 站点名称就像

NEW YORK PENN STATION, 
NEWARK PENN STATION,
NEWARK BROAD ST,
NEW PROVIDENCE
etc

当用户输入像 NEW YORK 这样的搜索输入时,我们会在结果中得到 NEW YORK PENN STATION,但当用户输入准确的 NEW YORK PENN STATION 在搜索输入中返回零个结果。

我的代码是 -

public ArrayList<String> getSimilarString(ArrayList<String> source, String querystr)
  {
      ArrayList<String> arResult = new ArrayList<String>();

        try
        {
            // 0. Specify the analyzer for tokenizing text.
            //    The same analyzer should be used for indexing and searching
            StandardAnalyzer analyzer = new StandardAnalyzer(Version.LUCENE_40);

            // 1. create the index
            Directory index = new RAMDirectory();

            IndexWriterConfig config = new IndexWriterConfig(Version.LUCENE_40, analyzer);

            IndexWriter w = new IndexWriter(index, config);

            for(int i = 0; i < source.size(); i++)
            {
                addDoc(w, source.get(i), "1933988" + (i + 1) + "z");
            }

            w.close();

            // 2. query
            // the "title" arg specifies the default field to use
            // when no field is explicitly specified in the query.
            Query q = new QueryParser(Version.LUCENE_40, "title", analyzer).parse(querystr + "*");

            // 3. search
            int hitsPerPage = 20;
            IndexReader reader = DirectoryReader.open(index);
            IndexSearcher searcher = new IndexSearcher(reader);
            TopScoreDocCollector collector = TopScoreDocCollector.create(hitsPerPage, true);
            searcher.search(q, collector);
            ScoreDoc[] hits = collector.topDocs().scoreDocs;

            // 4. Get results
            for(int i = 0; i < hits.length; ++i) 
            {
                  int docId = hits[i].doc;
                  Document d = searcher.doc(docId);
                  arResult.add(d.get("title"));
            }

            // reader can only be closed when there
            // is no need to access the documents any more.
            reader.close();

        }
        catch(Exception e)
        {
            System.out.println("Exception (LuceneAlgo.getSimilarString()) : " + e);
        }

        return arResult;

  }

  private static void addDoc(IndexWriter w, String title, String isbn) throws IOException 
  {
        Document doc = new Document();
        doc.add(new TextField("title", title, Field.Store.YES));

        // use a string field for isbn because we don't want it tokenized
        doc.add(new StringField("isbn", isbn, Field.Store.YES));
        w.addDocument(doc);
  }

在此代码中,source 是停止名称列表,query 是用户给定的搜索输入。

Lucene 算法是否适用于大字符串?

为什么 Lucene 算法不适用于 Exact String?

【问题讨论】:

  • 不确定答案,但您是否尝试过使用 Luke 查看您的索引? getopt.org/luke 可能是索引中的某些内容被篡改了。
  • 谢谢@Taylor,我会试试这个
  • “不工作”是什么意思?您的查询和返回结果是什么?
  • 您正在指定与停止名称相关的域,但代码中包含像“title”和“isbn”这样的字段名称,我真的很好奇..

标签: java lucene string-matching


【解决方案1】:

代替

1) Query q = new QueryParser(Version.LUCENE_40, "title", analyzer).parse(querystr + "*");

例如:“new york station”将被解析为“title:new title:york title:station”。 此查询将返回包含上述任何项的所有文档。

试试这个..

2) Query q = new QueryParser(Version.LUCENE_40, "title", analyzer).parse("+(" + querystr + ")");

Ex1:“new york”将被解析为“+(title:new title:york)”

上面的“+”表示“必须”在结果文档中出现该术语。 它将匹配包含“new york”和“new york station”的文档

Ex2:“纽约站”将被解析为 +(title:new title:york title:station)。 该查询将仅匹配“new york station”,而不仅仅是“new york”,因为 station 不存在。

请确保字段名称“标题”是您要查找的内容。

您的问题。

Lucene 算法是否适用于大字符串?

您必须定义什么是大字符串。你真的在寻找Phrase Search。一般来说,是的,Lucene 适用于大字符串。

为什么 Lucene 算法不适用于 Exact String?

因为解析 ("querystr" + "* ") 将生成单独的术语查询,并使用 OR 运算符连接它们。 例如:'new york*' 将被解析为:"title:new OR title:york*

如果您期待找到“纽约站”,则上述通配符查询不是您应该寻找的。这是因为您传入的 StandardAnalyser 在编制索引时会将纽约站标记(分解术语)为 3 个术语。

因此,查询“york*”将找到“york station”,只是因为它有“york”,而不是因为通配符,因为“york”不知道“station”,因为它们是不同的术语,即索引中的不同条目。

您真正需要的是PhraseQuery 用于查找确切的字符串,查询字符串应为“new york”带引号

【讨论】:

  • 感谢您的精彩解释,我会试试这个。
  • 嗨,我尝试了您建议的更改,但我没有得到预期的结果,即确切的字符串没有得到结果。我也尝试过 PhraseQuery,但没有运气(:
  • 1) 别担心,我会尽力解决您的问题。请将查询打印到系统输出并将其与查询字符串一起粘贴到此处。 2) 尝试使用 Luke 来检查您的索引。它是一个用于研究 Lucene Index 的 Java GUI 工具。您可以直接在工具中执行查询,这将加快您的调试速度。尝试 Luke(对于您的 Lucene 版本)并检查索引是否包含您正在寻找的术语。让我知道更新..
  • 嗨,Phani,我按照您的建议通过应用“+”让它工作了谢谢。抱歉周末回复晚了 :)
猜你喜欢
  • 2023-03-12
  • 1970-01-01
  • 2019-03-06
  • 2021-03-09
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多