【问题标题】:Java program to count Repeated lines in fileJava程序计算文件中的重复行
【发布时间】:2019-04-04 07:23:31
【问题描述】:
import java.io.BufferedReader;
import java.io.FileNotFoundException;
import java.io.FileReader;
import java.io.IOException;
import java.util.HashMap;
import java.util.Map;
import java.util.TreeMap;

public class work {

    public static void main(String[] args) throws FileNotFoundException, IOException {
        Map m1 = new HashMap();
        try (BufferedReader br = new BufferedReader(new FileReader("error.txt"))) {
            StringBuilder sb = new StringBuilder();
            String line = br.readLine();
            while (line != null) {
                String[] words = line.split(" ");//**This is where i was strucked**
                for (int i = 0; i < words.length; i++) {
                    if (m1.get(words[i]) == null) {
                        m1.put(words[i], 1);
                    } else {

                        int newValue = Integer.valueOf(String.valueOf(m1.get(words[i])));

                        newValue++;
                        m1.put(words[i], newValue);
                    }
                }
                sb.append(System.lineSeparator());
                line = br.readLine();
            }
        }
        Map<String, String> sorted = new TreeMap<String, String>(m1);
        for (Object key : sorted.keySet()) {
            System.out.println("Error : " + key + "Repeated " + m1.get(key) + " times.");
        }
    }

}

我有一个如下的文本文件,我想计算重复行的数量。我对如何拆分和计数感到震惊。谁能帮助我。

ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
    2018-10-29T12:01:00Z E! Error in plugin [inputs.openldap]: LDAP Result Code 32 "No Such Object": 
ERROR  [CompactionExecutor:21454] 2018-10-29 12:02:41,906 NoSpamLogger.java:91 - Maximum memory usage reached (125.000MiB), cannot allocate chunk of 1.000MiB
    2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
    2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
    2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
    2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml
    2018-09-20 14:08:14.571 [main] ERROR  org.apache.flink.yarn.YarnApplicationMasterRunner  -     -Dlogback.configurationFile=file:logback.xml

【问题讨论】:

  • 这可能会有所帮助 - stackoverflow.com/questions/46796021/…
  • 如果你只是计算重复的行,那你为什么需要拆分单词呢?
  • 那么如何计算重复行数
  • 创建一个Map&lt;String, Integer&gt;
  • 对于索引为 i 的每个句子,检查它是否与索引为 [0, i-1] 的任何句子重复。如果是,则跳过,否则通过与索引 [i+1, n-1] 处的每个句子的相似度来计算重复次数。最后添加所有这些计数。复杂度为 O(n^2*L),其中 L 是最长句子的长度。

标签: java


【解决方案1】:

如果您使用的是 Java8,请试试这个,目的是计算重复的行数(不是单词)

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Map;
import java.util.function.Function;
import java.util.stream.Collectors;

public class Work {
    public static void main(String[] args) throws IOException {
        Map<String, Long> dupes = Files.lines(Paths.get("/tmp/error.txt"))
                .collect(Collectors.groupingBy(Function.identity(), 
                     Collectors.counting()));

        // pretty print
        dupes.forEach((k, v)-> System.out.printf("(%d) times : %s ....%n", 
             v, k.substring(0,  Math.min(50, k.length()))));
    }
}

输出:

(2) times : 2018-09-20 14:08:14.571 [main] ERROR  org.apache.f ....
(8) times :     2018-10-29T12:01:00Z E! Error in plugin [input ....
(9) times : ERROR  [CompactionExecutor:21454] 2018-10-29 12:02 ....
(5) times :     2018-09-20 14:08:14.571 [main] ERROR  org.apac ....

【讨论】:

    【解决方案2】:

    Map&lt;String,Integer&gt; 可以与记录一起用作键和计数作为值。

        Map<String,Integer>  countMap= new HashMap<String,Integer>();
    
        try (
                BufferedReader  br= new BufferedReader(new FileReader(new File("D:\\error.txt")))
    
            ){
    
            String data="";
            while ((data=br.readLine())!=null) {
    
                if(countMap.containsKey(data)) {
                    countMap.put(data, countMap.get(data)+1);
                }else {
                    countMap.put(data, 1);
                }
    
            }
    
            countMap.forEach((k,v)->{System.out.println(k+" Occurs "+v+" times.");});
    
        } catch (IOException  e) {
            e.printStackTrace();
        }
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-12-05
      • 1970-01-01
      • 1970-01-01
      • 2016-04-26
      • 2015-06-11
      • 2020-10-09
      相关资源
      最近更新 更多