【问题标题】:Read and process 100 text files in c# in parallelc#中并行读取和处理100个文本文件
【发布时间】:2017-04-12 17:08:10
【问题描述】:

我有一个项目可以读取 100 个文本文件,其中包含 5000 个单词。

我将单词插入到列表中。我有第二个列表,其中包含英语停用词。我比较两个列表并从第一个列表中删除停用词。

运行应用程序需要 1 小时。我想并行化它。我该怎么做?

这是我的代码:

    private void button1_Click(object sender, EventArgs e)
    {

        List<string> listt1 = new List<string>();
        string line;

        for (int ii = 1; ii <= 49; ii++)
        {

            string d = ii.ToString();
            using (StreamReader reader = new StreamReader(@"D" + d.ToString() + ".txt"))

            while ((line = reader.ReadLine()) != null)
            {

                string[] words = line.Split(' ');
                for (int i = 0; i < words.Length; i++)
                {
                    listt1.Add(words[i].ToString());



                }
            }

            listt1 = listt1.ConvertAll(d1 => d1.ToLower());

            StreamReader reader2 = new StreamReader("stopword.txt");
            List<string> listt2 = new List<string>();
            string line2;
            while ((line2 = reader2.ReadLine()) != null)
            {
                string[] words2 = line2.Split('\n');
                for (int i = 0; i < words2.Length; i++)
                {
                    listt2.Add(words2[i]);

                }
                listt2 = listt2.ConvertAll(d1 => d1.ToLower());

            }

            for (int i = 0; i < listt1.Count(); i++)
            {
                for (int j = 0; j < listt2.Count(); j++)
                {
                    listt1.RemoveAll(d1 => d1.Equals(listt2[j]));

                }
            }
            listt1=listt1.Distinct().ToList();


            textBox1.Text = listt1.Count().ToString();
        }
    }
  }
}

【问题讨论】:

  • 如果运行时间那么长,就会出现问题。
  • 我用两个文件做的,列表计数是 1780 现在我用 49 个文件做了,它运行了 35 分钟
  • 运行两个文件需要多长时间?
  • 大约需要 3 分钟
  • 但文件字数不同,第一个有 1500 个字,第二个有 5000 个字

标签: c# multithreading file text parallel-processing


【解决方案1】:

我用你的代码解决了很多问题。我认为您不需要多线程:

    private void RemoveStopWords()
    {
        HashSet<string> stopWords = new HashSet<string>();

        using (var stopWordReader = new StreamReader("stopword.txt"))
        {
            string line2;
            while ((line2 = stopWordReader.ReadLine()) != null)
            {

                string[] words2 = line2.Split('\n');
                for (int i = 0; i < words2.Length; i++)
                {
                    stopWords.Add(words2[i].ToLower());
                }
            }
        }

        var fileWords = new HashSet<string>();

        for (int fileNumber = 1; fileNumber <= 49; fileNumber++)
        {               
            using (var reader = new StreamReader("D" + fileNumber.ToString() + ".txt"))
            {
                string line;
                while ((line = reader.ReadLine()) != null)
                {
                    foreach(var word in line.Split(' '))
                    {
                        fileWords.Add(word.ToLower());
                    }
                }
            }
        }

        fileWords.ExceptWith(stopWords);

        textBox1.Text = fileWords.Count().ToString();


    }

由于代码的结构方式,您多次阅读停用词列表,并不断添加到列表中并重新尝试删除相同的停用词。 HashSet 比 List 更符合您的需求,因为它已经处理了基于集合的操作和唯一性。

如果您仍然想使这个并行,您可以通过读取一次停用词列表并将其传递给将读取输入文件、删除停用词并返回结果列表的异步方法来做到这一点,那么您需要在异步调用返回后合并结果列表,但您最好在决定是否需要它之前进行测试,因为这比这段代码已经有更多的工作和复杂性。

【讨论】:

  • 我还建议 OP 了解算法复杂性的大 O 表示法:此处的 wiki 链接 en.wikipedia.org/wiki/Big_O_notation。很高兴知道循环操作和其他算法如何运行以更好地优化您的程序。
【解决方案2】:

如果我理解正确,您希望:

  1. 将文件中的所有单词读入列表中
  2. 从列表中删除所有“停用词”
  3. 再重复 99 个文件,只保存唯一的单词

如果正确,代码很简单:

// The list of words to delete ("stop words")
var stopWords = new List<string> { "remove", "these", "words" };

// The list of files to check - you can get this list in other ways
var filesToCheck = new List<string>
{
    @"f:\public\temp\temp1.txt",
    @"f:\public\temp\temp2.txt",
    @"f:\public\temp\temp3.txt"
};

// This list will contain all the unique words from all
// the files, except the ones in the "stopWords" list
var uniqueFilteredWords = new List<string>();

// Loop through all our files
foreach (var fileToCheck in filesToCheck)
{
    // Read all the file text into a varaible
    var fileText = File.ReadAllText(fileToCheck);

    // Split the text into distinct words (splitting on null 
    // splits on all whitespace) and ignore empty lines
    var fileWords = fileText.Split(null)
        .Where(line => !string.IsNullOrWhiteSpace(line))
        .Distinct();

    // Add all the words from the file, except the ones in 
    // your "stop list" and those that are already in the list
    uniqueFilteredWords.AddRange(fileWords.Except(stopWords)
        .Where(word => !uniqueFilteredWords.Contains(word)));
}

这可以浓缩成一行,没有显式循环:

// This list will contain all the unique words from all 
// the files, except the ones in the "stopWords" list
var uniqueFilteredWords = filesToCheck.SelectMany(fileToCheck =>
    File.ReadAllText(fileToCheck)
        .Split(null)
        .Where(word => !string.IsNullOrWhiteSpace(word) &&
                       !stopWords.Any(stopWord => stopWord.Equals(word, 
                           StringComparison.OrdinalIgnoreCase)))
        .Distinct());

此代码在不到一秒的时间内处理了超过 100 个文件,每个文件超过 12000 个单词(WAY 不到一秒...0.0001782 秒)

【讨论】:

  • 我试图在我的解决方案中保留他的代码结构,但我更喜欢你的。今天我了解了 .Split(null) 的作用(尽管没有什么可以原谅调用约定)。您还可以将 StringSplitOptions.None 传递给 split 方法,这样您就可以在 where 子句中删除 IsNullOrWhitespace 调用。
  • @wllmsaccnt 是的,但是在传递 null 时该选项不可用 - 您必须改为传递要分割的空白字符数组(这很好)
  • 寻求帮助
  • @RufusL 你教了我一个有用的晦涩的方法调用,所以我会报答:"".Split((char[])null, StringSplitOptions.RemoveEmptyEntries);你可以投 null,谁知道呢。
  • @wllmsaccnt 酷,我不知道!
【解决方案3】:

我在这里看到的一个可以帮助提高性能的问题是listt1.ConvertAll() 将在列表中以 O(n) 运行。您已经在循环将项目添加到列表中,为什么不在那里将它们转换为小写。另外为什么不将单词存储在哈希集中,这样您就可以在 O(1) 中查找和插入。您可以将停用词列表存储在哈希集中,当您阅读文本输入时,查看该词是否为停用词,如果不将其添加到哈希集中以输出用户。

【讨论】:

  • 我在 5 分钟前添加了一个答案,它完成了所有这些。他还多次阅读停用词列表,并从头开始不断添加要检查的单词列表,而不仅仅是他在该迭代中添加的那些。
  • 噢,伙计错过了时机,哈哈谢谢你用代码发布答案。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-01-24
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多