【问题标题】:how to read the text word by word如何逐字阅读文本
【发布时间】:2013-03-05 17:04:10
【问题描述】:

我正在处理 txt 或 htm 文件。目前我正在使用for循环逐个字符查找文档,但我需要逐个单词查找文本,然后逐个字符在单词内部查找。 我该怎么做?

for (int i = 0; i < text.Length; i++)
{}

【问题讨论】:

  • 您需要一种在文件中分隔单词的方法。空格可能会起作用,但我可以看到标点符号等问题......
  • 使用正则表达式匹配呈现单词的模式。然后按字符搜索匹配字符
  • htmlagilitypack.codeplex.com 是一个很好的与 .Net 一起使用的 HTML 解析器
  • @Alan 它对于文本文件可能工作得很好,但我认为可以安全地假设他的 .htm 文件包含 HTML 标记,用正则表达式解析会变得非常棘手。跨度>

标签: c# text streamreader


【解决方案1】:

一种简单的方法是使用不带参数的string.Split(由空格字符分割):

using (StreamReader sr = new StreamReader(path)) 
{
    while (sr.Peek() >= 0) 
    {
        string line = sr.ReadLine();
        string[] words = line.Split();
        foreach(string word in words)
        {
            foreach(Char c in word)
            {
                // ...
            }
        }
    }
}

我用StreamReader.ReadLine 阅读了整行。

要解析 HTML,我会使用像 HtmlAgilityPack 这样的强大库。

【讨论】:

    【解决方案2】:

    您可以在空格上拆分字符串,但您必须处理标点符号和 HTML 标记(您说您使用的是 txt 和 htm 文件)。

    string[] tokens = text.split(); // default for split() will split on white space
    foreach(string tok in tokens)
    {
        // process tok string here
    }
    

    【讨论】:

      【解决方案3】:

      这是我对StreamReader 的惰性扩展的实现。我们的想法是不要将整个文件加载到内存中,尤其是当您的文件是单行时。

      public static string ReadWord(this StreamReader stream, Encoding encoding)
      {
          string word = "";
          // read single character at a time building a word 
          // until reaching whitespace or (-1)
          while(stream.Read()
             .With(c => { // with each character . . .
                  // convert read bytes to char
                  var chr = encoding.GetChars(BitConverter.GetBytes(c)).First();
      
                  if (c == -1 || Char.IsWhiteSpace(chr))
                       return -1; //signal end of word
                  else
                       word = word + chr; //append the char to our word
      
                  return c;
          }) > -1);  // end while(stream.Read() if char returned is -1
          return word;
      }
      
      public static T With<T>(this T obj, Func<T,T> f)
      {
          return f(obj);
      }
      

      简单地使用:

      using (var s = File.OpenText(file))
      {
          while(!s.EndOfStream)
              s.ReadWord(Encoding.Default).ToCharArray().DoSomething();
      }
      

      【讨论】:

        【解决方案4】:

        使用text.Split(' ') 将其按空格拆分为单词数组,然后对其进行迭代。

        所以

        foreach(String word in text.Split(' '))
           foreach(Char c in word)
              Console.WriteLine(c);
        

        【讨论】:

          【解决方案5】:

          你可以在空格上分割:

          string[] words = text.split(' ')
          

          会给你一个单词数组,然后你可以遍历它们。

          foreach(string word in words)
          {
              word // do something with each word
          }
          

          【讨论】:

            【解决方案6】:

            我认为你可以使用拆分

                     var  words = reader.ReadToEnd().Split(' ');
            

            或使用

            foreach(String words in text.Split(' '))
               foreach(Char char in words )
            

            【讨论】:

              【解决方案7】:

              您可以使用HTMLAgilityPack 来get all the text from some HTML。如果你认为这太过分了,请看here。

              HtmlDocument doc = new HtmlDocument();
              doc.LoadHtml(text);
              
              foreach(HtmlNode node in doc.DocumentNode.SelectNodes("//text()"))
              {
                  var nodeText = node.InnerText;
              }
              

              然后您可以将每个节点的文本内容拆分为单词,一旦您定义了单词是什么。

              也许像this,

              using HtmlAgilityPack;
              
              static IEnumerable<string> WordsInHtml(string text)
              {
                  var splitter = new Regex(@"[^\p{L}]*\p{Z}[^\p{L}]*");
              
                  HtmlDocument doc = new HtmlDocument();
                  doc.LoadHtml(text);
              
                  foreach(HtmlNode node in doc.DocumentNode.SelectNodes("//text()"))
                  {
                      foreach(var word in splitter.Split(node.InnerText)
                      {
                          yield return word;
                      }
                  }
              }
              

              然后,检查每个单词中的字符

              foreach(var word in WordsInHtml(text))
              {
                  foreach(var c in word)
                  {
                      // a enumeration by word then char.
                  }
              }
              

              【讨论】:

                【解决方案8】:

                什么是正则表达式?

                using System;
                using System.Linq;
                using System.Text.RegularExpressions;
                
                namespace ConsoleApplication58
                {
                    class Program
                    {
                        static void Main()
                        {
                            string input =
                                @"I'm working with a txt or htm file. And currently I'm looking up the document char by char, using for loop, but I need to look up the text word by word, and then inside the word char by char. How can I do this?";
                            var list = from Match match in Regex.Matches(input, @"\b\S+\b")
                                       select match.Value; //Get IEnumerable of words
                            foreach (string s in list) 
                                Console.WriteLine(s); //doing something with it
                            Console.ReadKey();
                        }
                    }
                }
                

                它适用于任何分隔符,并且是最快的方法。

                【讨论】:

                  猜你喜欢
                  • 1970-01-01
                  • 2014-09-01
                  • 1970-01-01
                  • 1970-01-01
                  • 1970-01-01
                  • 2015-03-31
                  • 1970-01-01
                  • 1970-01-01
                  • 1970-01-01
                  相关资源
                  最近更新 更多