【问题标题】:regex taking over 1000 minutes to complete正则表达式需要超过 1000 分钟才能完成
【发布时间】:2018-12-15 14:53:09
【问题描述】:

下面的代码从名为 WORDS 的文件中读取单词列表,然后使用这些单词并在名为 CONTENT 的文件中查找它们,然后从 CONTENT 中删除这些单词并用###### 替换它们并创建一个名为 FINAL 的新文件 - words 文件有大约 16k 行单词,CONTENT 文件有大约 16k 行,总共大约 800 万个单词 - 当我运行它时,它需要 1000 多分钟才能完成,我最终放弃了。

有什么方法可以加快这个过程或使用更有效的方法吗? Words 中的单词以 \b 开头并以 \b 结尾 - 代码在我在较小的 CONTENT 文件上测试时确实有效

using System;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.Threading.Tasks;
using System.IO;
using System.Text.RegularExpressions;


namespace ConsoleApp10
{
    class Program
    {
        static void Main(string[] args)
        {
            string SAR_CONTACTS = @"C:\Users\root\Desktop\WORDS.csv";
            string SAR_CONTENT = @"C:\Users\root\Desktop\CONTENT.csv";
            string READ_SAR_CONTACTS;
            using (StreamReader streamReader = new StreamReader(SAR_CONTENT, Encoding.UTF8))
            READ_SAR_CONTACTS = streamReader.ReadToEnd();

            string SAR_CONTACTS_FILE = File.ReadAllText(SAR_CONTACTS);
            string SAR_CONTENT_FILE = SAR_CONTACTS_FILE.Replace("\r\n", "|");
            SAR_CONTENT_FILE = SAR_CONTENT_FILE.Remove(SAR_CONTENT_FILE.Length - 1);
            string SAR_CONTENT_CENSORED = Regex.Replace(READ_SAR_CONTACTS, SAR_CONTENT_FILE, "######", RegexOptions.IgnoreCase);
            File.WriteAllText(@"C:\Users\root\Desktop\FINAL.csv", SAR_CONTENT_CENSORED);
        }
    }
}

【问题讨论】:

  • 是的,用单词创建一个哈希集,然后只需从内容中一次读取一个单词,然后在创建 FINAL 的集合中查找它。
  • 看看这个看看对你有没有帮助codeproject.com/Articles/12383/…
  • 首先,请停止在C#中使用SHOUTING_SNAKE_FORM;它使您的代码难以阅读。在 C# 中,我们使用 camelCasedIdentifiers 表示本地人。其次,这是对正则表达式的一种非常糟糕的使用;它们不是为这项任务而设计的。如果您有多个替换项,请使用不同的搜索和替换机制。正则表达式替换旨在处理少量替换,而不是数百万。
  • InBetween - 单词列表远远超过 18k 单词,不确定如何实现
  • 人们会停止在这里标记 cmets。 Eric Lippert 给出了很好的建议。直截了当的建议并不粗鲁。

标签: c# regex multithreading visual-studio csv


【解决方案1】:

一般来说,我会简单地将 Regex 扔出窗口,因为对于如此庞大的文件,它会很快变得复杂。使用您的联系人文件,而不是\b,我可能会用一组分隔符替换它,例如£&%(如果联系人按该顺序使用相同的分隔符字符串,这将中断)。

这就是我写它的方式 - 请注意,就效率而言,这可能不是最有效的,但它会起作用。另请注意,我添加了 Replace 的 VB 版本,因此忽略大小写,因为 C# 版本没有此重载(您也可以编写扩展函数)。

using Microsoft.VisualBasic;
using System;
using System.Collections.Generic;
using System.IO;
using System.Linq;
using System.Text;
using System.Text.RegularExpressions;
using System.Threading.Tasks;

namespace ConsoleApp5
{
    class Program
    {
        static void Main(string[] args)
        {
            string contacts = @"contacts.csv";
            string content = @"content.csv";
            string[] delimiter = { "£&%" };
            string read_contents;

            using (StreamReader streamReader = new StreamReader(content, Encoding.UTF8))
                read_contents = streamReader.ReadToEnd();

            string sar_contacts = File.ReadAllText(contacts);
            List<string> contactsToReplace = sar_contacts.Split(delimiter, StringSplitOptions.RemoveEmptyEntries).ToList();

            int i = 0;
            foreach (var wordToCensor in contactsToReplace)
            {
                read_contents = Strings.Replace(read_contents, wordToCensor, "######", 1, -1, Constants.vbTextCompare);
                Console.WriteLine(++i); // so we know where we are
            }

            File.WriteAllText(@"filtered.csv", read_contents);
        }
    }
}

【讨论】:

  • 非常感谢您的详细回复 sushi7777 我将对此进行测试并确认它是否更快
  • 将所有 16k 行和 800 万字上传到内存中?
  • 为什么不呢?这是一种粗略的方式,但不会花费超过 1000 分钟,这是一个基本的编程示例,这就是我在这种情况下所要做的
  • 我刚刚在主文件上运行了代码,按照这个速度,它将在 9.4 小时内完成,哈哈
  • 比 16 小时快得多:D
猜你喜欢
  • 1970-01-01
  • 2018-05-20
  • 1970-01-01
  • 2010-10-19
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-11-15
  • 1970-01-01
相关资源
最近更新 更多