【问题标题】:Count occurences of phrases in lowercase sentences without punctuation计算没有标点符号的小写句子中短语的出现次数
【发布时间】:2021-10-10 10:36:31
【问题描述】:

示例输入:

  • 彼得警告彼得不要出海
  • 阿尔弗雷德应该去海边
  • 阿尔伯特警告皮特
  • 我告诉过你不要去迪士尼乐园

示例输出:

  • 3:去
  • 2:出海
  • 2:不去

可选:

  • 不计算低于 2 的出现次数。

我的方法

public void zähleHäufigkeitWorte(string[] arr, int n)
{
    bool[] visited = new bool[n];

    // Traverse through array elements and 
    // count frequencies 
    for (int i = 0; i < n; i++)
    {

        // Skip this element if already processed 
        if (visited[i] == true)
            continue;

        // Count frequency 
        int count = 1;
        for (int j = i + 1; j < n; j++)
        {
            if (arr[i] == arr[j])
            {
                visited[j] = true;
                count++;
            }
        }
        wort.Add(arr[i],count);
    }
}

如何扩展以计算短语的出现次数?

【问题讨论】:

  • 您可能需要一个Dictionary&lt;string, int&gt;,其中键是短语,值是出现计数器。
  • @PeterCsala 感谢您的编辑。

标签: c# count frequency find-occurrences


【解决方案1】:

这应该可以解决问题:

using System;
using System.Collections.Generic;

namespace SO68666232
{
    public class PhraseCounter
    {

        public static Dictionary<string, int> countWordAndPhraseFrequencies(string input)
        {
            var phraseDic = new Dictionary<string, int>();
            string[] tokens = input.Split(' ');
            int ntokens = tokens.Length;
            for (int start = 0; start < ntokens; start++)
            {
                for (int nwords = 1; nwords <= (ntokens- start) ; nwords++)
                {
                    ArraySegment<string> phraseTokens = new ArraySegment<string>(tokens, start, nwords);
                    string phrase = string.Join(" ", phraseTokens);
                    if (phraseDic.ContainsKey(phrase))
                    {
                        phraseDic[phrase]++;
                    }
                    else
                    {
                        phraseDic[phrase] = 1;
                    }
                }
            }
            // Remove elements with only 1 occurrence
            IEnumerable<string> keys = new List<string>(phraseDic.Keys);
            foreach (string key in keys)
            {
                if (phraseDic[key] <= 1)
                {
                    phraseDic.Remove(key);
                }
            }
            return phraseDic;
        }
    }
}

这可能不是最有效的方法,因为时间随着单词数的平方而增加,但它确实有效。

测试程序:

    class Program
    {
        const string sampleInput = "peter warns pete to not go to the sea alfred should go to the sea albert warns pete i told you to not go to disneyland";

        static void Main(string[] args)
        {
            var res = PhraseCounter.countWordAndPhraseFrequencies(sampleInput);
            foreach (string key in res.Keys)
            {
                Console.WriteLine("{0}: {1}", key, res[key]);
            }
            Console.WriteLine("Press any key to continue...");
            Console.ReadKey();
        }
    }

结果:

warns: 2
warns pete: 2
pete: 2
to: 5
to not: 2
to not go: 2
to not go to: 2
not: 2
not go: 2
not go to: 2
go: 3
go to: 3
go to the: 2
go to the sea: 2
to the: 2
to the sea: 2
the: 2
the sea: 2
sea: 2

【讨论】:

  • 你好。我在string phrase = string.Join(" ", phraseTokens); 收到System.OutOfMemoryException。
  • 您使用的是此处给出的测试字符串还是更大的输入?你输入的大小是多少?该算法基本上列出了每个可能的短语,因此如果输入中有 n 个单词,则将有 O(n^2) 个平均长度为 O(n) 的可能短语,因此内存需求随着 O(n^ 3)。该算法可以很容易地进行修改,以遵守 10 个单词的最大短语长度,然后内存需求将显着降低。
  • 另一种选择是修改算法以立即计算短语的出现次数,并且仅在 > 1 时保存,而不是累积每个短语并在最后进行过滤。
猜你喜欢
  • 2022-11-03
  • 1970-01-01
  • 2016-09-01
  • 2021-02-25
  • 1970-01-01
  • 2022-11-28
  • 1970-01-01
  • 2021-06-11
  • 1970-01-01
相关资源
最近更新 更多