【问题标题】:Fortran Dot Product performance: Increasing cpu time consumption with diminishing array sizeFortran 点积性能:随着数组大小的减小而增加 CPU 时间消耗
【发布时间】:2021-04-12 19:22:55
【问题描述】:

这更多是关于自相关函数和计算机性能的理论问题。

https://en.wikipedia.org/wiki/Autocorrelation#Estimation

(抱歉不确定如何在堆栈溢出时键入方程式)

自相关函数大多只是一个数组移动了一个时间(索引)的点积。

因此,自相关使用以下的点积:

原始数组从其初始索引 (0) 到大小 - 相关时间索引

x

以及从 index(correlation_time) 到 size 的原始数组

下面是这个点积总和的简单图片。 Another Picture

[1 , 2, 3, 4, 5, 6, 7, 8, 9, 10]
x    |  |  |  |  |  |  |  |  |
    [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]

[1 , 2, 3, 4, 5, 6, 7, 8, 9, 10]
x                   |  |  |  | 
                   [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]

据此,我相信当您长时间计算自相关时,点积应该会明显变小并且应该更容易计算。然而,下面的 fortran 代码却另有说明。

随着相关时间增加,因此要进行点积的数组更小,计算时间随着相关时间幅度的增加而变大!

我到底错过了什么?

acorr.f90

PROGRAM acorr
    real:: a,b,c,d, sum
    integer:: i,j, jsize,beginning, rate, end
    real, dimension(2000000):: Jx, corr
    integer:: skip_lines = 4
    call system_clock(beginning, rate)
   
   !Reading file here
    open(10, file='data.log', status='old')
    do i = 1,skip_lines
        read(10,*)
    end do
    do i = 1, 2000000
        read(10,*) Jx(i)
    end do
    call system_clock(end)
    print *, "elapsed time for reading: ", real(end - beginning) / real(rate)
    close(10)
    jsize = size(Jx)
    print *, "Size of Jx: ", jsize
    call system_clock(end)
    !End of reading file

    !begin dot product for autocorrelation
    print *, "Loop Start Time: ", real(end - beginning) / real(rate)
    do i =1,jsize
        a = dot_product(Jx(i:jsize),Jx(1:jsize-(i-1)))
        if(i == 1) then
            call system_clock(end)
            print *, "correlation time magnitude 1e0 elapsed time: ", real(end - beginning) / real(rate)

        else if(i == 10) then
            call system_clock(end)
            print *, "correlation time magnitude 1e1 elapsed time: ", real(end - beginning) / real(rate)


        else if(i == 100) then
            call system_clock(end)
            print *, "correlation time magnitude 1e2 elapsed time: ", real(end - beginning) / real(rate)

        else if(i == 1000) then
            call system_clock(end)
            print *, "correlation time magnitude 1e3 elapsed time: ", real(end - beginning) / real(rate)
        else if(i == 10000) then
            call system_clock(end)
            print *, "correlation time magnitude 1e4 elapsed time: ", real(end - beginning) / real(rate)
        else if(i == 100000) then
            call system_clock(end)
            print *, "correlation time magnitude 1e5 elapsed time: ", real(end - beginning) / real(rate)
        else if(i == 1000000) then
            call system_clock(end)
            print *, "correlation time magnitude 1e6 elapsed time: ", real(end - beginning) / real(rate)


        end if 

    end do
    call system_clock(end)
    print *, "elapsed time: ", real(end - beginning) / real(rate)
END PROGRAM

输出

elapsed time for reading:    4.67100000    
 Size of Jx:      2000000
 Loop Start Time:    4.67100000    
 correlation time magnitude 1e0 elapsed time:    4.67299986    
 correlation time magnitude 1e1 elapsed time:    4.69500017    
 correlation time magnitude 1e2 elapsed time:    4.90100002    
 correlation time magnitude 1e3 elapsed time:    6.93699980    
 correlation time magnitude 1e4 elapsed time:    27.3729992    
 correlation time magnitude 1e5 elapsed time:    227.809006 
 correlation time magnitude 1e6 elapsed time:    1704.59399    
 elapsed time:    2249.08105 

注意:我有一个包含 2,000,0000 个时间点的数据文件。

编译: gfortran -o acorr.exe acorr.f90

系统:Ubuntu Linux 20.04

编辑 我只是把点积线改成只做整个数组的点积

a = dot_product(Jx,Jx)

结果仍然相同,并且随着循环索引的增加,它需要更长的时间。

编辑 2

看起来我没有正确理解我的输出。 我在循环中添加了以下内容:

    a = dot_product(Jx(i:jsize),Jx(1:jsize-(i-1)))
    a = 0
    call system_clock(end)
    write(20,*) real(end - end1) / real(rate)

看起来每个循环的输出只有 1 毫秒。这大约是完整向量的单个 dot_product 所需的时间。所以我相信它正在按预期工作。原始输出应该以对数尺度进行解释,因为显然从 10,000 到 20,000 比从 100 到 200 需要更多的迭代。

【问题讨论】:

  • 你可能错过了很多,但我们不知道,因为你没有告诉计算环境。什么操作系统?什么编译器标志?您是否为点积使用了选项而不是使用 BLAS?此外,使用数组部分作为实际参数可能会导致编译器使用临时变量。
  • 最后更新了注释。即使数组部分使用临时变量不会改变这样一个事实,即随着数组变小,点积应该更快。

标签: performance fortran dot-product autocorrelation


【解决方案1】:

看起来我没有正确理解我的输出。我在循环中添加了以下内容:

a = dot_product(Jx(i:jsize),Jx(1:jsize-(i-1)))
a = 0
call system_clock(end)
write(20,*) real(end - end1) / real(rate)

看起来每个循环的输出只有 1 毫秒。这大约是完整向量的单个 dot_product 所需的时间。所以我相信它正在按预期工作。原始输出应该以对数尺度解释,因为显然从 10,000 到 20,000 比从 100 到 200 需要更多的迭代。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-06-07
    • 2021-09-27
    • 2015-06-01
    • 2020-06-23
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多