【问题标题】:how to improve foreach loop performance in C#如何提高 C# 中的 foreach 循环性能
【发布时间】:2020-07-30 10:58:18
【问题描述】:
class Student
{
    public int ID { get; set; }
    public string Name { get; set; }
    public string Email { get; set; }
}

class Program
{
    private static object _lockObj = new object();

    static void Main(string[] args)
    {

        List<int> collection = Enumerable.Range(1, 100000).ToList();


         
        List<Student> students = new List<Student>(100000);
        var options = new ParallelOptions()
        {
            MaxDegreeOfParallelism = 1
        };
        var sp = System.Diagnostics.Stopwatch.StartNew();
        sp.Start();
        Parallel.ForEach(collection, options, action =>
        {
            lock (_lockObj)
            {
                 var dt = collection.FirstOrDefault(x => x == action);
                 if (dt > 0)
                 {
                    Student student = new Student();
                    student.ID = dt;
                    student.Name = "Zoyeb";
                    student.Email = "ShaikhZoyeb@Gmail.com";

                    students.Add(student);
                    Console.WriteLine(@"value of i = {0}, thread = {1}", 
                    action,Thread.CurrentThread.ManagedThreadId);
                }
            }
        });
        sp.Stop();

        double data = Convert.ToDouble(sp.ElapsedMilliseconds / 1000);
        Console.WriteLine(data);
         
    }
}

我想尽快循环 100000 条记录 我尝试了 foreach 循环,但循环遍历 100000 条记录并不是很好,然后在我尝试实现 Parallel.ForEach() 以提高我的性能之后,在实际场景中我将收集 Id,我需要查找集合是否 id 退出如果退出然后添加。 性能达到最佳状态 当我评论条件时执行大约需要 3 秒,当我取消注释条件时大约需要 24 秒所以我的问题是有什么方法可以通过在集合中查找 id 来提高我的性能

         //var dt = collection.FirstOrDefault(x => x == action);
         //if (dt > 0)
         //{
            Student student = new Student();
            student.ID = 1;
            student.Name = "Zoyeb";
            student.Email = "ShaikhZoyeb@Gmail.com";

            students.Add(student);
            Console.WriteLine(@"value of i = {0}, thread = {1}", 
            action,Thread.CurrentThread.ManagedThreadId);
        //}

【问题讨论】:

  • parallel.foreach 中有一个锁,这意味着每次迭代都将按顺序运行,而不是并行运行。基本上,您采用了 parallel.foreach 并将其转换回一个过于复杂的普通 foreach。您将不得不重组您的代码,以便您可以移除该锁定。
  • 你能给我看一个@Lasse V. Karlsen 的例子吗?
  • 很遗憾没有,目前仅在 iPhone 上。你需要使用可以被多个线程同时操作和使用的集合类型。
  • 处理性能问题的最佳方法是首先分析/测量,然后更改代码。除非您知道瓶颈在哪里,否则不要开始进行更改。
  • 是的,我完全同意你@Lasse V. Karlsen

标签: c#


【解决方案1】:

您的原始代码在Parallel.ForEach 中执行lock。这实质上是采用并行代码并强制它串行运行。

在我的机器上需要 40 秒。

真的相当于这样做:

    foreach (var action in collection)
    {
            var dt = collection.FirstOrDefault(x => x == action);
            if (dt > 0)
            {
                Student student = new Student();
                student.ID = dt;
                student.Name = "Zoyeb";
                student.Email = "ShaikhZoyeb@Gmail.com";

                students.Add(student);
            }
    }

这也需要 40 秒。

但是,如果你这样做:

    foreach (var action in collection)
    {
        Student student = new Student();
        student.ID = action;
        student.Name = "Zoyeb";
        student.Email = "ShaikhZoyeb@Gmail.com";

        students.Add(student);
    }

运行需要 1 毫秒。它大约快 40,000 倍。

在这种情况下,您可以通过迭代您的集合一次来获得更快的循环,而不是以嵌套方式并且不使用Parallel.ForEach。


我很抱歉错过了关于 id 不存在的部分。

试试这个:

    HashSet<int> hashSet = new HashSet<int>(collection);

    List<Student> students = new List<Student>(100000);

    var sp = System.Diagnostics.Stopwatch.StartNew();
    sp.Start();
    foreach (var action in collection)
    {
        if (hashSet.Contains(action))
        {
            Student student = new Student();
            student.ID = action;
            student.Name = "Zoyeb";
            student.Email = "ShaikhZoyeb@Gmail.com";

            students.Add(student);
        }
    }
    sp.Stop();

运行时间为 3 毫秒。

另一种方法是像这样使用join:

    foreach (var action in
        from c in collection
        join dt in collection on c equals dt
        select dt)
    {
        Student student = new Student();
        student.ID = action;
        student.Name = "Zoyeb";
        student.Email = "ShaikhZoyeb@Gmail.com";

        students.Add(student);
    }

运行时间为 25 毫秒。

【讨论】:

  • 我知道如果我删除 FirstOrDefault() 条件会快得多,但是在实际情况下我们有选择助手,用户将选择行和进程的范围,这就是我必须在集合中查找的原因,我很清楚有问题提到我知道如果 FirstOrDefault() 被删除,那么它会非常快速地工作
【解决方案2】:

问题 1

您在并行 foreach 中使用锁而不是并发集合。强制并行 foreach 等待,访问锁,因此执行将是一个接一个。

将您的列表更改为 ConcurrentBag,并从 ParallelForEach 中删除 lock

// using System.Collections.Concurrent; // at the top
var students = new ConcurrentBag<Student>()

问题 2

FirstOrDefault() 如果您想按 Id 选择,性能不是很好。使用字典。由于字典执行哈希匹配,因此它比 FirstOrDefault 快得多。见thisquestoin。

将您的收藏更改为字典:

var collection = Enumerable.Range(1, 100000)
    .ToDictionary(x=> x);

将循环中的访问更改为:

if(collection.TryGetValue(action, out var dt))
{
  //....
}

问题 3

秒表不是基准测试工具。请使用Benchmark.Net 或其他库。

【讨论】:

  • 这是一个糟糕的解决方案。使用ConcurrentBag 是在现有大量黑客之上的又一黑客。
  • 我花了大约 5 秒来迭代,非常感谢@Preben Huybrechts
  • 你是什么意思@Enigmativity 我没明白?
  • @Enigmativity,请解释一下?正如我read Concurrent bag 的性能一样。
  • @ZoyebShaikh - 看看我的回答。它在 1 毫秒内运行。
猜你喜欢
  • 1970-01-01
  • 2018-02-24
  • 2023-03-09
  • 1970-01-01
  • 2021-06-19
  • 1970-01-01
  • 2020-08-25
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多