【发布时间】:2016-10-07 23:46:05
【问题描述】:
在我的 main.cpp 中有一段摘录:
Ptr<FastFeatureDetector> fastDetector = FastFeatureDetector::create(80, true);
while (true) {
Mat image = // get grayscale image 1280x720
timer.start();
detector->detect(image, keypoints);
myfile << "FAST\t" << timer.end() << endl; // timer.end() is how many seconds elapsed since last timer.start()
keypoints.clear();
timer.start();
for (int i = 3; i < image.rows - 3; i++)
{
for (int j = 3; j < image.cols - 3; j++)
{
if (inspectPoint(image.data, image.cols, i, j)) {
// this block is never entered
KeyPoint keypoint(i, j, 3);
keypoints.push_back(keypoint);
}
}
}
myfile << "Custom\t" << timer.end() << endl;
myfile << endl;
myfile.flush();
...
}
myfile 说:
FAST 0.000515495
Custom 0.00221361
FAST 0.000485697
Custom 0.00217653
FAST 0.000490001
Custom 0.00219044
FAST 0.000484373
Custom 0.00216329
FAST 0.000561184
Custom 0.00233214
所以人们会认为inspectPoint() 是一个实际上在做某事的函数。
bool inspectPoint(const uchar* img, int cols, int i, int j) {
uchar p = img[i * cols + j];
uchar pt = img[(i - 3)*cols + j];
uchar pr = img[i*cols + j + 3];
uchar pb = img[(i + 3)*cols + j];
uchar pl = img[i*cols + j - 3];
return cols < pt - pr + pb - pl + i; // just random check so that the optimizer doesn't skip any calculations
}
我使用的是 Visual Studio 2013,优化设置为“完全优化 (/Ox)”。
据我所知,FAST 算法会遍历所有像素?我想它实际上不可能比函数inspectPoint()更快地处理每个像素。
FAST 检测器为何如此之快?或者说,为什么嵌套循环这么慢?
【问题讨论】:
-
通过快速浏览源代码,fastFeatureDetector 中似乎对 SSE 和 OpenCL 进行了广泛的优化:github.com/Itseez/opencv/blob/master/modules/features2d/src/…
-
那么,如果我理解正确的话,通过为特定的 CPU 编写代码,您可以获得超过 4 倍的加速吗?或者,FAST 检测器可能会尽可能使用 GPU?这两个似乎都不可信,但我不是专家......
-
SSE 和 OpenCL 并不特定于任何 CPU。 SSE 利用 CPU 对多条数据同时执行单个指令(计算)的能力。因此,根据 CPU 的架构,这可以将速度提高 2 倍或远远超过 4 倍。 OpenCL 可以利用 GPU,它还可以显着提升某些图像处理操作的性能。
-
BHawk,如果您可以复制您的评论作为答案,我会接受。也许你也可以看看相关的问题:stackoverflow.com/questions/39920813/…
标签: c++ opencv feature-detection