【问题标题】:4 slices of bytearray to compare using multithreading使用多线程比较 4 片字节数组
【发布时间】:2012-09-03 19:12:35
【问题描述】:

我被项目的第二阶段困住了: 将一个字节[] 分成 4 个切片(以加载 QuadCore I5 CPU),然后在每个切片上,在每个内核上启动一个线程(比较任务)。

原因是试图加速两个相同大小的字节数组之间的比较 这个怎么穿?

    [DllImport("msvcrt.dll", CallingConvention = CallingConvention.Cdecl)]
    static extern int memcmp(byte[] b1, byte[] b2, long count);




class ArrayView<T> : IEnumerable<T> 
    {
        private readonly T[] array;
        private readonly int offset, count;
        public ArrayView(T[] array, int offset, int count)
        {
            this.array = array; this.offset = offset; this.count = count;
        }
        public int Length { get { return count; } }
        public T this[int index] {
            get { if (index < 0 || index >= this.count)
                throw new IndexOutOfRangeException();
            else
                return this.array[offset + index]; }
            set { if (index < 0 || index >= this.count)
                throw new IndexOutOfRangeException();
            else
                this.array[offset + index] = value; }
        }
        public IEnumerator<T> GetEnumerator()
        {
            for (int i = offset; i < offset + count; i++)
                yield return array[i];
        } 
        IEnumerator IEnumerable.GetEnumerator()
        {
            IEnumerator<T> enumerator = this.GetEnumerator();
            while (enumerator.MoveNext())
            {
                yield return enumerator.Current;
            }
        }
    } 




    public void CopmarArrSlice()
    {

    byte[] LoadedArr = File.ReadAllBytes("testFileCompare2Scr.bmp");

    int LoddArLn = OrgArr.Length;
    int range =  (LoddArLn / 4) - LoddAremainder;
    int divisionremain = LoddArLn - (range * 4);

    ArrayView<byte> LddP1 = new ArrayView<byte>(OrgArr, 0, range);
    ArrayView<byte> LddP2 = new ArrayView<byte>(OrgArr, p1.Length, range);
    ArrayView<byte> LddP3 = new ArrayView<byte>(OrgArr, (p1.Length + p2.Length), range);
    ArrayView<byte> LddP4 = new ArrayView<byte>(OrgArr, (p1.Length + p2.Length + p3.Length), range + divisionremain);


        if (AreEqual(LddP1, CapturedP1)) ....Do Somthing

    }


    public bool AreEqual(byte[] a, byte[] b)
    {
        if (a == b)
           return true;
        if (a == null || b == null)
            return false;
        if (a.Length != b.Length)
            return false;
        return memcmp(a, b, a.Length) == 0;
    }


    CopmarArrSlice();

在这种情况下,如何使用 AreEqual(使用 memcmp)将其与使用 4 个线程/Parallelism 进行比较,在每个 CpuCore 上进行计算

【问题讨论】:

  • 我完全希望多线程会使这个任务变慢

标签: c# multithreading arraylist parallel-processing bytearray


【解决方案1】:

我编写了一个尽可能利用多核的函数,但它似乎受到 p/invoke 调用的严重性能影响。我认为这个版本只有在测试非常大的数组时才有意义。

static unsafe class NativeParallel
{
    [DllImport("msvcrt.dll", CallingConvention = CallingConvention.Cdecl)]
    static extern int memcmp(byte* b1, byte* b2, int count);

    public static bool AreEqual(byte[] a, byte[] b)
    {
        // The obvious optimizations
        if (a == b)
            return true;
        if (a == null || b == null)
            return false;
        if (a.Length != b.Length)
            return false;

        int quarter = a.Length / 4;
        int r0 = 0, r1 = 0, r2 = 0, r3 = 0;
        Parallel.Invoke(
            () =>  {
                fixed (byte* ap = &a[0])
                fixed (byte* bp = &b[0])
                    r0 = memcmp(ap, bp, quarter);                        
            },
            () => {
                fixed (byte* ap = &a[quarter])
                fixed (byte* bp = &b[quarter])
                    r1 = memcmp(ap, bp, quarter);
            },
            () => {
                fixed (byte* ap = &a[quarter * 2])
                fixed (byte* bp = &b[quarter * 2])
                    r2 = memcmp(ap, bp, quarter);
            },
            () => {
                fixed (byte* ap = &a[quarter * 3])
                fixed (byte* bp = &b[quarter * 3])
                    r3 = memcmp(ap, bp, a.Length - (quarter * 3));
            }
        );
        return r0 + r1 + r2 + r3 == 0;
    }
}

在大多数情况下,它实际上比优化的安全版本要慢。

static class SafeParallel
{
    public static bool AreEqual(byte[] a, byte[] b)
    {
        if (a == b)
            return true;
        if (a == null || b == null)
            return false;
        if (a.Length != b.Length)
            return false;

        bool b1 = false;
        bool b2 = false;
        bool b3 = false;
        bool b4 = false;
        int quarter = a.Length / 4;
        Parallel.Invoke(
            () => b1 = AreEqual(a, b, 0, quarter),
            () => b2 = AreEqual(a, b, quarter, quarter),
            () => b3 = AreEqual(a, b, quarter * 2, quarter),
            () => b4 = AreEqual(a, b, quarter * 3, a.Length)
        );
        return b1 && b2 && b3 && b4;
    }

    static bool AreEqual(byte[] a, byte[] b, int start, int length)
    {
        var len = length / 8;
        if (len > 0)
        {
            for (int i = start; i < len; i += 8)
            {
                if (BitConverter.ToInt64(a, i) != BitConverter.ToInt64(b, i))
                    return false;
            }
        }
        var remainder = length % 8;
        if (remainder > 0)
        {
            for (int i = length - remainder; i < length; i++)
            {
                if (a[i] != b[i])
                    return false;
            }
        }
        return true;
    }
}

【讨论】:

  • 只是让你知道我在你 ChaosBAcomp 之后命名了一个 byteArray 比较类(:形成上一篇文章,现在对它们进行基准测试,再次感谢
  • ok 将代码(第一个)复制到我的项目中,错误很少 1):---错误6 'MyScrCaptu_Comp.MyForm1.memcmp(byte[], byte[ ], long)' ...MyForm1.cs 246 22 MyScrCaptu_Comp ....(#2)Error 5 Cannot convert lambda expression to type 'System.Threading.Tasks.ParallelOptions 因为它不是委托类型 243 13 ...(#3)...Error 8 参数 2 : 无法从 'byte*' 转换为 'byte[]' 246 33
  • @LoneXcoder - 我认为您需要将memcmp 的签名切换为static extern int memcmp(byte* b1, byte* b2, int count); 。
  • 抱歉没有复制你的,因为我有一个(:试试看
  • .ok 上次它是原生的(pinvoke memcmp)但安全,而不是并行,现在它不安全和并行,去 Chek,基准测试......请不要屏住呼吸;我会花一些时间.. 20-40 分钟才能再次通过它们。回来..我会发布结果。
【解决方案2】:

我认为您不必以单线程和经典 c# 方式拆分字节

foreach(byte currentArr in LoadedArr)
{
    if (AreEqyal(currentArr, CapturedP1))
           ....Do Somthing
}

但是要通过将工作负载分配给多个线程来处理每个字节,您必须使用以下语法;

// max your threads count in my case 16,
int[] sums = new int[16];// optional, just to know the workload
public void ProcessMyByte(byte current)
{
    if (AreEqyal(current, CapturedP1))
           ....Do Somthing

       // optional just to know what thread is in
        sums[Thread.CurrentThread.ManagedThreadId]++;// increment the number of iterations done by the thread who did this elementary process
}

.... Main()...
{
....

byte[] LoadedArr = File.ReadAllBytes("testFileCompare2Scr.bmp");
Parallel.ForEach(LoadedArr, ProcessMyByte);

...
}

因此并行性将由您代为管理,并且会更好,因为当线程空闲时,它会执行下一个任务,而不是您将其拆分为 4 个,每个线程必须处理 Length/4。

【讨论】:

  • 我不会浪费时间拆分数组,图像以 Byte[] 的形式存储在一个文件夹中,然后我将所有 byte[] 加载到内存中,将它们拆分为切片,然后转储原始数组并使用只有部分数据进行比较,我想如果我只比较 4 个切片中的 2|3 个,如果它们真的不一样,我仍然会有一个错误。你是否仍然认为不拆分我不明白你发布的代码(第一个)以及它是如何不必要的 - 切片比较,切片与切片与完整数组对比完整数组,除非你知道你在说什么并且有一个我的知识理解有问题
  • 去看看这个...所以我的ProcessMyByte() 做的是计算相等性,然后 不管 true ||假它“增加 therad”Parallel.ForEach(),所以你可以通过引用 where(&under what search term)来打开它对于它,或者如果你真的有耐心一步一步地描述它,谢谢你的帮助Séddik!
  • 看看这个链接books.google.com/…
  • 抱歉英语不好,我很着急,不增加线程而是计算它的迭代次数,例如我已经尝试使用长度为 20 的 byte[] 执行parallel.foreach,线程号 2 执行 6 次操作,N°9 执行 13 次,N°11 执行其余操作。
  • 这本书很棒,几乎涵盖了所有内容
猜你喜欢
  • 1970-01-01
  • 2014-11-02
  • 2010-11-26
  • 1970-01-01
  • 1970-01-01
  • 2021-09-02
  • 1970-01-01
  • 2023-04-04
  • 2016-12-20
相关资源
最近更新 更多