【问题标题】:n-gram character based similarity measure基于 n-gram 字符的相似性度量
【发布时间】:2023-04-06 02:32:01
【问题描述】:

我已经使用代码从单词中提取了二元组:

Scanner a = new Scanner(file1);
PrintWriter pw1= new PrintWriter(file2);  
    while (a.hasNext()) {
       String gram=a.next();
       pw1.println(gram);
       int line;
       line = gram.length(); 
       for(int x=0;x<=line-2;x++){
         pw1.println(gram.substring(x, x+2));        
       }
    }
    pw1.close();
}
catch(FileNotFoundException e) {
  System.err.format("FileNotExist");`
}

例如,“student”的二元组是“st”、“tu”、“ud”、“de”、“en”、“nt”。

但我需要找到相似度计算。

我必须计算这些拆分后的克之间的相似度值。

【问题讨论】:

  • 相似度值是什么意思?
  • 您到底想问什么?首先让你自己清楚,然后让你的代码易于理解,以便我们可以帮助你。
  • As,即两个二元组的 Jaccard?您的 sn-p 中缺少 try{...

标签: java


【解决方案1】:

嗯,你没有很好地解释你的问题,但这是我的看法。

首先,您的代码到处都是,即使是这么小的程序,任何人都很难阅读任何内容,我会对其进行编辑以使其可读,但我不确定是否允许。

无论如何,bigrams = 成对的字母、单词或音节(根据 Google 的说法)。并且您正在寻找它的相似度值?

嗯,我对它做了一点研究,看来你需要的算法就是这个。


现在,让我们开始充实这一点。 OP,如果我误解了你的问题,请纠正我。您正在寻找单词之间的相似性,方法是将它们分解为双元组并找到它们各自的相似性值,是吗?如果是这样,让我们​​在开始使用之前分解这个公式,因为它肯定是你需要的。


现在,我们有 2 个单词,FRANCE 和 FRENCH。如果将它们分解为二元组,我们希望找到它们的相似值。

对于法国,我们有 {FR, RA, AN, NC, CE}
对于法语,我们有 {FR, RE, EN, NC, CH}

France 和 French 用于表示第一个等式中的 s1 和 s2。接下来,我们取他们在两个二元组中的匹配项。我的意思是,在这两个词中都可以找到哪些二元组或字母对?在这种情况下,答案是 FR 和 NC。

因为我们找到了 2 对,所以顶部的值变成 4,因为公式表明,2 乘以匹配的二元组数。所以我们在顶部有 4 个,仅此而已。

现在,通过将您要比较的每个单词可以组成的二元组数相加来解决下半部分的问题,即 5 表示 FRANCE,5 表示 FRENCH。所以分母是 10

那么现在,我们有什么?我们有 4/10 或 0.4。 THAT 是我们的相似值,也是您在制作程序时需要找到的值。


让我们尝试另一个例子,让我们把它根植于我们的脑海中,比如说

s1 = "歌利亚"
s2 = "守门员"

所以,使用二元组,我们得到了字符串数组...

{"GO","OL","LI","IA","AT","TH"} {"GO","OA","AL","LI","IE"}

现在,匹配的数量。这两个词中有多少个匹配的二元组?答案 - 2、GO 和 LI

所以,分子应该有

2 x {2 个匹配项} = 4

现在,分母,我们有 6 个用于歌利亚的二元组和 5 个用于守门员的二元组。请记住,我们必须按照原始公式添加这 2 个值,因此我们的分母将是 11。

那么,这让我们何去何从?

S(GOLIATH, GOALIE) = 4/11 ~ .364


我在这个链接下找到了公式(基本上是我刚刚学到的所有东西),这真的很容易理解。

http://www.catalysoft.com/articles/StrikeAMatch.html

我将编辑此评论,因为我需要一段时间才能为您的课程提出一个方法,但只是为了快速回复,如果您正在寻找有关如何操作的更多帮助,链接是开始的好地方。

编辑****

好的,刚刚为它构建了方法,就在这里。

public class BiGram
{

/*

here are your imports

import java.util.Scanner;
import java.io.File;
import java.io.PrintWriter;
import java.io.FileNotFoundException;

*/
//you'll have to forgive the lack of order or common sense, I threw it 
//together fast I could cuz it sounded like you were in a rush

   public String[][] bigramizedWords = new String[2][100];

   public String[] words = new String[2];

   public File file1 = new File("file1.txt");
   public File file2 = new File("file2.txt");

   public int tracker = 0;
   public double matches = 0;
   public double denominator = 0; //This will hold the sum of the bigrams of the 2 words

   public double results;

   public Scanner a;
   public PrintWriter pw1;


   public BiGram()
   {

      initialize();
      bigramize();

      results = matches/denominator;

      pw1.println("\n\nThe Bigram Similarity value between " + words[0] + " and " + words[1] + " is " + results  + ".");


      pw1.close();


   }

   public static void main(String[] args)
   {

      BiGram b = new BiGram();


   }

   public void initialize()
   {

      try
      {

         a = new Scanner(file1);
         pw1 = new PrintWriter(file2);

         while (a.hasNext()) 
         {

            //System.out.println("Enter 2 words delimited by a space to calculate their similarity values based off of bigrams.");
            //^^^ I was going to use the above line, but I see you are using File and PrintWriter, 
            //I assume you will have the files yourself with the words to be compared

            String gram  = a.next();

            //pw1.println(gram);    -----you had this originally, we don't need this
            int line = gram.length(); 

            for(int x=0;x<=line-2;x++)
            {

               bigramizedWords[tracker][x] = gram.substring(x, x+2);
               pw1.println(gram.substring(x, x+2) + "");

            }

            pw1.println("");

            words[tracker] = gram;

            tracker++;

         }


      }

      catch(FileNotFoundException e) 
      {
         System.err.format("FileNotExist");
      }
   }

   public void bigramize()
   {

      denominator = (words[0].length() - 1) + (words[1].length() - 1); 
      //^^ Let me explain this, basically, for every word you have, let's say BABY and BALL,
      //the denominator is gonna be the sum of number of bigrams. In this case, BABY has {BA,AB,BY} or 3
      //bigrams, same for BALL, {BA,AL,LL} or 3. And the length of the word BABY is 4 letters, same 
      //with Ball. So basically, just subtract their respective lengths by 1 and add them together, and 
      //you get the number of bigrams combined from both words, or 6


      for(int k = 0; k < bigramizedWords[0].length; k++)
      {

         if(bigramizedWords[0][k] != null)
         {


            for(int i = 0; i < bigramizedWords[1].length; i++)
            {

            ///////////////////////////////////////////

               if(bigramizedWords[1][i] != null)
               {

                  if(bigramizedWords[0][k].equals(bigramizedWords[1][i]))
                  {

                     matches++;

                  }

               }

            }

         }

      }

      matches*=2;




      }

}

【讨论】:

  • 是的,但是 OP 并没有给我们太多的工作,我认为有总比没有好,特别是因为 OP 似乎很匆忙。我即将发布我想出的解决这个问题的代码。我会在完成一些错误后发布它,也许您可​​以检查我是否以正确的方式进行操作?了解,我在不到 2 小时前就学会了这一切,我不知道自己在做什么,我只是按照指示行事。听起来你比我更了解这一点。
  • 绝对正确!我当然不是在批评你!我只是提供更多信息以获得更完整的答案。
  • 哦,好吧,那我的错,误读了。不管怎样,马上就要发布代码了。
猜你喜欢
  • 2011-05-01
  • 2015-01-22
  • 1970-01-01
  • 2016-05-02
  • 1970-01-01
  • 2019-05-01
  • 2017-03-15
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多