【发布时间】:2018-12-15 14:53:09
【问题描述】:
下面的代码从名为 WORDS 的文件中读取单词列表,然后使用这些单词并在名为 CONTENT 的文件中查找它们,然后从 CONTENT 中删除这些单词并用###### 替换它们并创建一个名为 FINAL 的新文件 - words 文件有大约 16k 行单词,CONTENT 文件有大约 16k 行,总共大约 800 万个单词 - 当我运行它时,它需要 1000 多分钟才能完成,我最终放弃了。
有什么方法可以加快这个过程或使用更有效的方法吗? Words 中的单词以 \b 开头并以 \b 结尾 - 代码在我在较小的 CONTENT 文件上测试时确实有效
using System;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.Threading.Tasks;
using System.IO;
using System.Text.RegularExpressions;
namespace ConsoleApp10
{
class Program
{
static void Main(string[] args)
{
string SAR_CONTACTS = @"C:\Users\root\Desktop\WORDS.csv";
string SAR_CONTENT = @"C:\Users\root\Desktop\CONTENT.csv";
string READ_SAR_CONTACTS;
using (StreamReader streamReader = new StreamReader(SAR_CONTENT, Encoding.UTF8))
READ_SAR_CONTACTS = streamReader.ReadToEnd();
string SAR_CONTACTS_FILE = File.ReadAllText(SAR_CONTACTS);
string SAR_CONTENT_FILE = SAR_CONTACTS_FILE.Replace("\r\n", "|");
SAR_CONTENT_FILE = SAR_CONTENT_FILE.Remove(SAR_CONTENT_FILE.Length - 1);
string SAR_CONTENT_CENSORED = Regex.Replace(READ_SAR_CONTACTS, SAR_CONTENT_FILE, "######", RegexOptions.IgnoreCase);
File.WriteAllText(@"C:\Users\root\Desktop\FINAL.csv", SAR_CONTENT_CENSORED);
}
}
}
【问题讨论】:
-
是的,用单词创建一个哈希集,然后只需从内容中一次读取一个单词,然后在创建 FINAL 的集合中查找它。
-
看看这个看看对你有没有帮助codeproject.com/Articles/12383/…
-
首先,请停止在C#中使用SHOUTING_SNAKE_FORM;它使您的代码难以阅读。在 C# 中,我们使用
camelCasedIdentifiers表示本地人。其次,这是对正则表达式的一种非常糟糕的使用;它们不是为这项任务而设计的。如果您有多个替换项,请使用不同的搜索和替换机制。正则表达式替换旨在处理少量替换,而不是数百万。 -
InBetween - 单词列表远远超过 18k 单词,不确定如何实现
-
人们会停止在这里标记 cmets。 Eric Lippert 给出了很好的建议。直截了当的建议并不粗鲁。
标签: c# regex multithreading visual-studio csv