【问题标题】:How to check if file contains strings, which will double repeat sign?如何检查文件是否包含字符串,这将重复符号?
【发布时间】:2014-05-29 16:22:30
【问题描述】:

我想检查包含一些字符串的文件,用# 分隔是否包含双重复符号。例子: 我有一个这样的文件:

1234#224859#123567

我正在阅读此文件并将用# 分隔的字符串放入数组中。 我想找出哪些字符串有一个相邻重复的数字(在本例中为 224859)并返回在此字符串中重复的第一个数字的位置?

这是我目前所拥有的:

    ArrayList list = new ArrayList();
    OpenFileDialog openFile1 = new OpenFileDialog();

        int size = -1;
        DialogResult dr = openFile1.ShowDialog();
        string file = openFile1.FileName;
        try
        {
            string text = File.ReadAllText(file);
            size = text.Length;
            string temp = "";

            for (int i = 0; i < text.Length; i++)
            {
                if (text[i] != '#')
                {
                    temp += text[i].ToString();
                }
                else
                {
                    list.Add(temp);
                    temp = "";
                }
            }
        }
        catch (IOException)
        {
        }
        string all_values = "";
        foreach (Object obj in list)
        {
            all_values += obj.ToString() + " => ";

            Console.WriteLine(" => ", obj);
        }
        textBox1.Text = (all_values);

【问题讨论】:

  • 也许只有我一个人,但我不知道你想做什么。 “我想找出哪些字符串包含双重复符号(在本例中为 224859)”这是否意味着具有任何双字符的字符串?
  • 一些示例输入/输出在这里会非常好。似乎是 String.Split() 的一个很好的用例,但如果没有更多信息就很难判断。
  • 我的意思是究竟是哪个数字,两个相同的数字相邻!希望现在很清楚:D
  • 您只需要22 或224859?
  • @SriramSakthivel 我相信他需要0,重复的索引。

标签: c# file openfiledialog


【解决方案1】:

这个正则表达式应该可以解决问题。

var subject = "1234#224859#123567";
foreach(var item in subject.Split('#'))
{
    var regex = new Regex(@"(?<grp>\d)\k<grp>");
    var match =regex.Match(item);
    if(match.Success)
    {
         Console.WriteLine("Index : {0}, Item:{1}", match.Index, item);
        //prints Index : 0, Item:224859
    }
}

【讨论】:

    【解决方案2】:

    这是一种比 Sriram 更程序化的方法,但主要好处是记住您的结果,以便以后在您的程序中使用它们。

    基本上,字符串是根据# 分隔符拆分的,它返回一个string[],其中包含每个数字。然后,对于每个字符串,您遍历字符并检查i 处的当前字符是否与i + 1 处的下一个字符匹配。如果是这样,重复数字最早出现在i,所以i被记住,我们跳出处理chars的循环。

    由于int 是不可为空的类型,我决定使用-1 来表示在字符串中未找到匹配项。

    Dictionary<string, int> results = new Dictionary<string, int>();
    string text = "1234#224859#123567#11#4322#43#155";
    string[] list = text.Split('#');
    foreach (string s in list)
    {
        int tempResult = -1;
        for (int i = 0; i < s.Length - 1; i++)
        {
            if(s.ElementAt(i) == s.ElementAt(i + 1))
            {
                tempResult = i;
                break;
            }
        }
        results.Add(s, tempResult);
    }
    
    foreach (KeyValuePair<string, int> pair in results) 
    {
        Console.WriteLine(pair.Key + ": " + pair.Value);
    }
    

    输出:

    1234:-1

    224859: 0

    123567:-1

    11:0

    4322:2

    43:-1

    155: 1

    【讨论】:

    • 我比我更喜欢你的解决方案。字典和 ElementAt 非常优雅,Ryan。
    【解决方案3】:

    这是另一个有效的正则表达式

          int indexof = -1;
            String input = "3492883#32280948093284#990303294";
            string[] numbers = input.Split('#');
            foreach(string n in numbers)
            {
                Match m=Regex.Match(n, @"(\d)\1+");
                if (m.Success)
                {
                    indexof = m.Index;
                }
    
            }
    

    【讨论】:

    • 基本相同,不是另一个正则表达式。我使用了命名组,而你使用的索引就是不同的地方。
    【解决方案4】:

    这能满足您的需求吗?

    string text = File.ReadAllText(file);
    
    string[] list = text.Split(new char[] { '#' });
    

    然后,将字符串分开后:

            foreach (string s in list)
            {
                int pos = HasDoubleCharacter(s);
                if (pos > -1)
                {
                    // do something
                }
            }
    
        private static int HasDoubleCharacter(string text)
        {
            int pos = 0;
            char[] c3 = text.ToCharArray();
            char lastChar = (char)0;
            foreach (char c in c3)
            {
                if (lastChar == c)
                    return pos;
                lastChar = c;
                pos++;
            }
            return -1;
        }
    

    或者你只是在寻找原文中所有双打的位置列表。如果是这样(并且您不需要单独处理各种字符串,您可以试试这个:

        private static List<int> FindAllDoublePositions(string text)
        {
            List<int> positions = new List<int>();
            char[] ca = text.ToCharArray();
            char lastChar = (char)0;
            for (int pos = 0; pos < ca.Length; pos++)
            {
                if (Char.IsNumber(ca[pos]) && lastChar == ca[pos])
                    positions.Add(pos);
                lastChar = ca[pos];
            }
            return positions;
        }
    

    【讨论】:

    • 显然不是(虽然这也是我最初的想法)
    • 好吧,我还没有得到他想要做的其他事情。这个应该可以的。不像使用正则表达式那么优雅,但我非常喜欢简单易懂。
    • 此外,从示例中看起来文件中的所有字符都是数字,但如果您预期非数字,只需添加此测试: if (Char.IsNumber(c) && lastChar = = c)
    • 您也可以考虑如何处理非标准案例。如果你有两个分数标记在一起怎么办? (123#345##4556)。你想如何处理“三元组”?如果一个字符串是 #12333456" 你只想知道第一个位置 (2) 还是你看到两个双精度数 (2,3)?
    【解决方案5】:

    如果您正在寻找特定的字符串模式,Regex 很可能是您最好的朋友:

    string text = "1234#224859#123567asdashjehqwjk4234#244859#123567asdhajksdhqjkw1434#244859#123567";
    
    var results = Regex.Matches(text, @"\d{4}#(?<Value>\d{6})#\d{4}");
    var allValues = "";
    
    foreach (Match result in results)
    {
        allValues = result.Groups["Value"].Value + " => ";
        Console.WriteLine(" => ", result.Value);
    }
    

    【讨论】:

      猜你喜欢
      • 2019-06-06
      • 2012-04-21
      • 2017-02-10
      • 2020-09-20
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多