【问题标题】:Viola Jones AdaBoost running out of memory before even startsViola Jones AdaBoost 在启动之前就内存不足
【发布时间】:2012-11-25 04:22:56
【问题描述】:

我正在实施 Viola Jones 人脸检测算法。我在算法的 AdaBoost 学习部分的第一部分遇到问题。

原论文指出

弱分类器选择算法如下进行。对于每个特征,示例都根据特征值进行排序。

我目前正在处理一个相对较小的训练集,其中包含 2000 个正图像和 1000 个负图像。该论文描述了拥有多达 10,000 个数据集。

AdaBoost 的主要目的是减少 24x24 窗口中的特征数量,总计 160,000+。该算法对这些特征起作用并选择最好的。

论文描述了对于每个特征,它在每个图像上计算其值,然后根据值对它们进行排序。这意味着我需要为每个特征创建一个容器并存储所有样本的值。

我的问题是我的程序在仅评估了 10,000 个功能(其中只有 6%)后内存不足。所有容器的总大小最终将达到 160,000*3000,即数十亿。我应该如何在不耗尽内存的情况下实现这个算法?我增加了堆大小,它使我从 3% 增加到了 6%,我认为增加它不会起作用。

论文暗示在整个算法中都需要这些排序值,所以我不能在每个特征之后丢弃它们。

这是我目前的代码

public static List<WeakClassifier> train(List<Image> positiveSamples, List<Image> negativeSamples, List<Feature> allFeatures, int T) {
    List<WeakClassifier> solution = new LinkedList<WeakClassifier>();

    // Initialize Weights for each sample, whether positive or negative
    float[] positiveWeights = new float[positiveSamples.size()];
    float[] negativeWeights = new float[negativeSamples.size()];

    float initialPositiveWeight = 0.5f / positiveWeights.length;
    float initialNegativeWeight = 0.5f / negativeWeights.length;

    for (int i = 0; i < positiveWeights.length; ++i) {
        positiveWeights[i] = initialPositiveWeight;
    }
    for (int i = 0; i < negativeWeights.length; ++i) {
        negativeWeights[i] = initialNegativeWeight;
    }

    // Each feature's value for each image
    List<List<FeatureValue>> featureValues = new LinkedList<List<FeatureValue>>();

    // For each feature get the values for each image, and sort them based off the value
    for (Feature feature : allFeatures) {
        List<FeatureValue> thisFeaturesValues = new LinkedList<FeatureValue>();

        int index = 0;
        for (Image positive : positiveSamples) {
            int value = positive.applyFeature(feature);
            thisFeaturesValues.add(new FeatureValue(index, value, true));
            ++index;
        }
        index = 0;
        for (Image negative : negativeSamples) {
            int value = negative.applyFeature(feature);
            thisFeaturesValues.add(new FeatureValue(index, value, false));
            ++index;
        }

        Collections.sort(thisFeaturesValues);

        // Add this feature to the list
        featureValues.add(thisFeaturesValues);
        ++currentFeature;
    }

    ... rest of code

【问题讨论】:

  • 原始论文说:“鉴于检测器的基本分辨率为 24x24,矩形特征的详尽集合相当大,45,396”。不是160,000。 160,000 是如何获得的?
  • 您也不应该明确存储所有功能。一次只是为所有训练补丁中的一个特征提取的值。在 boosting 算法的每一步,您都可以评估每个特征的有用性,选择最好的一个,并将其添加到您的强分类器中。您永远不需要同时获得内存中所有图像的所有特征的结果。
  • 5 个特征中的每一个都具有大约 45,396 个可能的位置/大小。该论文确实提到了 160,000 个总特征。 2x2 特征(对角线区域)的可能性较小,这就是它的总和。
  • 好的。我的答案仍然有效。

标签: computer-vision face-detection adaboost viola-jones


【解决方案1】:

这应该是选择弱分类器之一的伪代码:

normalize the per-example weights  // one float per example

for feature j from 1 to 45,396:
  // Training a weak classifier based on feature j.
  - Extract the feature's response from each training image (1 float per example)
  // This threshold selection and error computation is where sorting the examples
  // by feature response comes in.
  - Choose a threshold to best separate the positive from negative examples
  - Record the threshold and weighted error for this weak classifier

choose the best feature j and threshold (lowest error)

update the per-example weights

您无需在任何地方存储数十亿个特征。只需在每次迭代中即时提取特征响应。您正在使用积分图像,因此提取速度很快。那是主要的内存瓶颈,它并不多,每个图像中的每个像素只有一个整数......基本上与图像所需的存储量相同。

即使您只是计算所有图像的所有特征响应并将它们全部保存,这样您就不必每次迭代都这样做,但仍然只是:

  • 45396 * 3000 * 4 字节 =~ 520 MB,或者如果您确信有 160000 个可能的功能,
  • 160000 * 3000 * 4 字节 =~ 1.78 GB,或者如果您使用 10000 个训练图像,
  • 160000 * 10000 * 4 字节 =~ 5.96 GB

基本上,即使您确实存储了所有特征值,也不应该耗尽内存。

【讨论】:

  • 这是否像论文描述的那样被另一个循环 t = 1, ... T 包围?您如何更新示例权重?您确定您选择的特征是否正确分类每个示例?当您选择一个特征然后想要寻找另一个特征时,您是否会从特征列表中删除先前选择的特征? (所以你不要再选它了)。我问这些是因为我的印象是你一开始就存储了所有的值,然后你就找到了这些特征。由于我的记忆问题,我想这并不聪明。我想我的内存用完了,因为我存储的是一个对象而不是数字。
  • 是的,这只是内部循环。这是完成T 次。更新示例权重的公式在论文中(他们的伪代码中的第 4 步。我可以重复一遍,但也许这对于单独的问题是最好的)。当您选择一个特征(给定迭代中的最佳特征)时,您通过存储特征 id、阈值和权重(三个浮点数)将其添加到强分类器中。然后,移动到下一个迭代并找到新的最佳特征。不过,这些似乎是关于算法的问题,而不是内存使用情况,所以最好单独提出一个问题。
  • 如果您发布后续问题,您可以在此处为我(和其他用户)粘贴他们的链接。
猜你喜欢
  • 2012-04-27
  • 2018-06-05
  • 1970-01-01
  • 2012-04-30
  • 2017-06-13
  • 2014-06-03
  • 2013-12-12
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多