【发布时间】:2015-03-15 15:53:41
【问题描述】:
我正在试用 Delphi XE7 Update 1 的并行编程功能。
我创建了一个简单的TParallel.For 循环,它基本上做了一些虚假操作来打发时间。
我在 AWS 实例 (c4.8xlarge) 上的 36 个 vCPU 上启动了该程序,以尝试了解并行编程的好处。
当我第一次启动程序并执行TParallel.For 循环时,我看到了显着的收益(尽管承认比我预期的 36 个 vCPU 少很多):
Parallel matches: 23077072 in 242ms
Single Threaded matches: 23077072 in 2314ms
如果我没有关闭程序并在不久之后(例如,立即或大约 10-20 秒后)在 36 个 vCPU 机器上再次运行传递,并行传递会恶化很多:
Parallel matches: 23077169 in 2322ms
Single Threaded matches: 23077169 in 2316ms
如果我不关闭程序并等待几分钟(不是几秒钟,而是几分钟),然后再次运行传递,我会再次获得第一次启动程序时得到的结果(10 倍改进响应时间)。
在 36 个 vCPU 的机器上,启动程序后的第一遍总是更快,所以这种效果似乎只在程序中第二次调用 TParallel.For 时发生。
这是我正在运行的示例代码:
unit ParallelTests;
interface
uses
Winapi.Windows, Winapi.Messages, System.SysUtils, System.Variants, System.Classes, Vcl.Graphics,
System.Threading, System.SyncObjs, System.Diagnostics,
Vcl.Controls, Vcl.Forms, Vcl.Dialogs, Vcl.StdCtrls;
type
TForm1 = class(TForm)
Button1: TButton;
Memo1: TMemo;
SingleThreadCheckBox: TCheckBox;
ParallelCheckBox: TCheckBox;
UnitsEdit: TEdit;
Label1: TLabel;
procedure Button1Click(Sender: TObject);
private
{ Private declarations }
public
{ Public declarations }
end;
var
Form1: TForm1;
implementation
{$R *.dfm}
procedure TForm1.Button1Click(Sender: TObject);
var
matches: integer;
i,j: integer;
sw: TStopWatch;
maxItems: integer;
referenceStr: string;
begin
sw := TStopWatch.Create;
maxItems := 5000;
Randomize;
SetLength(referenceStr,120000); for i := 1 to 120000 do referenceStr[i] := Chr(Ord('a') + Random(26));
if ParallelCheckBox.Checked then begin
matches := 0;
sw.Reset;
sw.Start;
TParallel.For(1, MaxItems,
procedure (Value: Integer)
var
index: integer;
found: integer;
begin
found := 0;
for index := 1 to length(referenceStr) do begin
if (((Value mod 26) + ord('a')) = ord(referenceStr[index])) then begin
inc(found);
end;
end;
TInterlocked.Add(matches, found);
end);
sw.Stop;
Memo1.Lines.Add('Parallel matches: ' + IntToStr(matches) + ' in ' + IntToStr(sw.ElapsedMilliseconds) + 'ms');
end;
if SingleThreadCheckBox.Checked then begin
matches := 0;
sw.Reset;
sw.Start;
for i := 1 to MaxItems do begin
for j := 1 to length(referenceStr) do begin
if (((i mod 26) + ord('a')) = ord(referenceStr[j])) then begin
inc(matches);
end;
end;
end;
sw.Stop;
Memo1.Lines.Add('Single Threaded matches: ' + IntToStr(Matches) + ' in ' + IntToStr(sw.ElapsedMilliseconds) + 'ms');
end;
end;
end.
这是否按设计工作?我发现这篇文章 (http://delphiaball.co.uk/tag/parallel-programming/) 建议我让库来决定线程池,但是如果我必须从请求到请求等待几分钟以便更快地处理请求,我看不到使用并行编程的意义。
我是否遗漏了关于应该如何使用 TParallel.For 循环的任何内容?
请注意,我无法在 AWS m3.large 实例(根据 AWS 为 2 个 vCPU)上重现此问题。在那种情况下,我总是会得到轻微的改善,并且在不久之后的TParallel.For 的后续调用中我没有得到更糟糕的结果。
Parallel matches: 23077054 in 2057ms
Single Threaded matches: 23077054 in 2900ms
因此,当有许多可用内核(36)时,似乎会出现这种效果,这很遗憾,因为并行编程的全部意义在于从许多内核中受益。我想知道这是否是库错误,因为核心数较多,或者在这种情况下核心数不是 2 的幂。
更新:使用不同 vCPU 的各种实例对其进行测试后 在 AWS 中很重要,这似乎是行为:
- 36 个 vCPU (c4.8xlarge)。您必须在后续调用普通 TParallel 调用之间等待几分钟(这使得它无法用于 生产)
- 32 个 vCPU (c3.8xlarge)。您必须在后续调用普通 TParallel 调用之间等待几分钟(这使得它无法用于 生产)
- 16 个 vCPU (c3.4xlarge)。您必须等待第二次。如果负载低但响应时间仍然很重要,它可能是可用的
- 8 个 vCPU (c3.2xlarge)。它似乎可以正常工作
- 4 个 vCPU (c3.xlarge)。它似乎可以正常工作
- 2 个 vCPU (m3.large)。它似乎可以正常工作
【问题讨论】:
-
@Pep 如果您认为库有问题,请使用另一个库编写代码并进行比较。我怀疑图书馆是问题所在。
-
我对它进行了更多测试。看来,至少对于 AWS,Parallel 库在 vCPU > 8 时会出现一些问题。vCPU = 16 时,它比 vCPU = 32 或 36 时要好得多,但它仍然存在问题。可能 TParallel.For 调用已针对最多 8 个虚拟内核(台式机)的系统进行了微调。我会用我的发现更新问题。
-
这不太可能。我不知道你为什么会这么想。不要猜测。在 OTL 下运行等效代码,看看会发生什么。
-
所以,我认为该库不会针对特定数量的内核进行优化。但我认为图书馆可能是问题的根源。如果是这样的话,与 OTL 的比较会给出一个好主意。我必须说的是,新的 RTL 并行库完全是垃圾。这里有无数的帖子表明它实施得非常糟糕。我怀疑我能否让自己使用它。我向你推荐 OTL。
-
我认为很明显,在某些时候您遇到了并行库中的错误,导致代码串行执行。这肯定不是设计使然。肯定不是因为调整错误。这肯定要归结为粗制滥造的实施。坦率地说,Embarcadero 在生成正确的线程代码方面有着糟糕的记录。在
TMonitor的惨败之后,还有人相信他们吗?
标签: delphi parallel-processing rtl-ppl