【问题标题】:Does this algorithm actually work? Sum of subsets backtracking algorithm这个算法真的有效吗?子集和回溯算法
【发布时间】:2019-11-10 05:35:26
【问题描述】:

我想知道这种回溯算法是否真的有效。

在教科书Foundations of Algorithms,第5th版中,定义如下:

算法 5.4:子集和问题的回溯算法

问题:给定n个正整数(权重)和一个正整数W, 确定总和为 W 的整数的所有组合。

输入:正整数n,排序(非降序)数组 从 1 到 n 索引的正整数 w,以及一个正整数 W.

输出:总和为 W 的所有整数组合。

void sum_of_subsets(index i, 
                    int weight, int total)  {
  if (promising(i))
     if (weight == W)
        cout << include[1] through include [i];
     else {
        include[i + 1] = "yes";               // Include w[i + 1].
        sum_of_subsets(i + 1, weight + w[i + 1], total - w[i + 1]);
        include[i + 1] = "no";                // Do not include w[i + 1].
        sum_of_subsets(i + 1, weight, total - w[i + 1]);
     }
}

bool promising (index i);  {
  return (weight + total >= W) && (weight == W || weight + w[i + 1] <= W);
}

按照我们通常的约定,n、w、W、 和 include 不是 输入到我们的例程中。如果这些变量是全局定义的,则 对 sum_of_subsets 的顶级调用如下:

sum_of_subsets(0, 0, total);

在第 5 章的结尾,exercise 13 问:

  1. 对子集和问题使用回溯算法(算法 5.4) 找到以下数字的所有组合,总和为 W = 52:

    w1 = 2    w2 = 10    w3 = 13    w4 = 17    w5 = 22    w6 = 42

我已经实现了这个精确的算法,考虑到从 1 开始的数组,它只是不起作用......

 void sos(int i, int weight, int total) {
    int yes = 1;
    int no = 0;

    if (promising(i, weight, total)) {
        if (weight == W) {
            for (int j = 0; j < arraySize; j++) {
                std::cout << include[j] << " ";
            }
            std::cout << "\n";
        }
        else if(i < arraySize) {
            include[i+1] = yes;
            sos(i + 1, weight + w[i+1], total - w[i+1]);
            include[i+1] = no;
            sos(i + 1, weight, total - w[i+1]);
        }
    }
}


int promising(int i,  int weight, int total) {
    return (weight + total >= W) && (weight == W || weight + w[i+1] <= W);
}   

我相信问题出在这里:

sos(i + 1, weight, total - w[i+1]);
sum_of_subsets(i+1, weight, total-w[i+1]);

当你到达这条线时,你没有正确回溯。

是否有人能够识别此算法的问题或实际对其进行编码以使其工作?

【问题讨论】:

  • 这不是最低可重现代码。数组w 和include 是如何定义的?
  • 是我自己还是这太复杂了?有前途的功能看起来像代码混淆比赛的参赛者......

标签: c++ algorithm backtracking


【解决方案1】:

我个人认为该算法有问题。没有边界检查,它使用了很多全局变量,并且假设数组从 1 开始索引。我认为您不能逐字复制它。它是实际实现的伪代码。在 C++ 中,数组总是从 0 开始。所以当你尝试做 include[i+1] 并且你只检查 i &lt; arraySize 时,你可能会遇到问题。

该算法还假设您有一个名为 total 的全局变量,它由函数 promising 使用。

我对代码进行了一些修改,将其放在一个类中,并进行了一些简化:

class Solution
{
private:
    vector<int> w;
    vector<int> include;

public:
    Solution(vector<int> weights) : w(std::move(weights)),
        include(w.size(), 0) {}

    void sos(int i, int weight, int total) {
        int yes = 1;
        int no = 0;
        int arraySize = include.size();

        if (weight == total) {
            for (int j = 0; j < arraySize; j++) {
                if (include[j]) {
                    std::cout << w[j] << " ";
                }
            }
            std::cout << "\n";
        }
        else if (i < arraySize)
        {
            include[i] = yes;
            //Include this weight
            sos(i + 1, weight + w[i], total);
            include[i] = no;
            //Exclude this weight
            sos(i + 1, weight, total);
        }
    }
};

int main()
{   
    Solution solution({ 2, 10, 13, 17, 22, 42 });
    solution.sos(0, 0, 52);
    //prints:    10 42
    //           13 17 22
}

【讨论】:

  • sos(i + 1, weight, total);是的,这就是我的想法,因为 sos(i + 1, weight, total - w[i+1]);就该算法应该如何工作而言,这没有任何意义。
【解决方案2】:

是的,正如其他人指出的那样,您偶然发现了基于 1 的数组索引。

除此之外,我认为您应该要求作者退还您为这本书支付的部分费用,因为他的代码逻辑过于复杂。

避免遇到边界问题的一个好方法是不使用 C++(期待这个大声笑的反对票)。

只有 3 个案例需要测试:

  • 候选值大于剩余值。 (破获)
  • 候选值正是剩余的值。
  • 候选值小于剩余值。

promising 函数试图表达这一点,然后在主函数 sos 中再次重新测试该函数的结果。

但它可能看起来像这样简单:

search :: [Int] -> Int -> [Int] -> [[Int]]
search (x1:xs) t path 
    | x1 > t = []
    | x1 == t = [x1 : path]
    | x1 < t = search xs (t-x1) (x1 : path) ++ search xs t path
search [] 0 path = [path]
search [] _ _ = []


items = [2, 10, 13, 17, 22, 42] :: [Int]
target = 52 :: Int

search items target []
-- [[42,10],[22,17,13]]

现在,在编写 C++ 代码时实现类似的安全网绝非不可能。但这需要决心和有意识地决定你愿意应对什么,什么不愿意。而且你需要愿意多输入几行代码才能完成 Haskell 的 10 行代码。

首先,我被原始 C++ 代码中索引和范围检查的所有复杂性所困扰。如果我们查看我们的 Haskell 代码(适用于列表), 确认我们根本不需要随机访问。我们只看剩余项目的开始。我们将一个值附加到路径(在 Haskell 中我们附加到前面是因为速度),最终我们将找到的组合附加到结果集中。考虑到这一点,对索引的困扰有点过头了。

其次,我更喜欢搜索功能的外观 - 显示 3 个关键测试,周围没有任何噪音。我的 C++ 版本应该努力做到漂亮。

此外,全局变量是如此 1980 - 我们不会有那个。将这些“全局变量”塞入一个类以稍微隐藏它们是 1995 年的事。我们也不会这样做。

就在这里! “更安全”的 C++ 实现。而且更漂亮……嗯……你们中的一些人可能不同意;)

#include <cstdint>
#include <vector>
#include <iostream>

using Items_t = std::vector<int32_t>;
using Result_t = std::vector<Items_t>;

// The C++ way of saying: deriving(Show)
template <class T>
std::ostream& operator <<(std::ostream& os, const std::vector<T>& value) 
{
    bool first = true;
    os << "[";
    for( const auto item : value) 
    {
        if(first) 
        {
            os << item;
            first = false;
        }
        else
        {
            os << "," << item;
        }
    }
    os << "]";
    return os;
}

// So we can do easy switch statement instead of chain of ifs.
enum class Comp : int8_t 
{   LT = -1
,   EQ = 0
,   GT = 1
};

static inline 
auto compI32( int32_t left, int32_t right ) -> Comp
{
    if(left == right) return Comp::EQ;
    if(left < right) return Comp::LT;
    return Comp::GT;
}

// So we can avoid index insanity and out of bounds problems.
template <class T>
struct VecRange
{
    using Iter_t = typename std::vector<T>::const_iterator;
    Iter_t current;
    Iter_t end;
    VecRange(const std::vector<T>& v)
        : current{v.cbegin()}
        , end{v.cend()}
    {}
    VecRange(Iter_t cur, Iter_t fin)
        : current{cur}
        , end{fin}
    {}
    static bool exhausted (const VecRange<T>&);
    static VecRange<T> next(const VecRange<T>&);
};

template <class T>
bool VecRange<T>::exhausted(const VecRange<T>& range)
{
    return range.current == range.end;
}

template <class T>
VecRange<T> VecRange<T>::next(const VecRange<T>& range)
{
    if(range.current != range.end)
        return VecRange<T>( range.current + 1, range.end );
    return range;   
}

using ItemsRange = VecRange<Items_t::value_type>;


static void search( const ItemsRange items, int32_t target, Items_t path, Result_t& result)
{
    if(ItemsRange::exhausted(items))
    {
        if(0 == target)
        {
            result.push_back(path);
        }
        return;
    }

    switch(compI32(*items.current,target))
    {
        case Comp::GT: 
            return;
        case Comp::EQ:
            {
                path.push_back(*items.current);
                result.push_back(path);
            }
            return;
        case Comp::LT:
            {
                auto path1 = path; // hope this makes a real copy...
                path1.push_back(*items.current);
                search(ItemsRange::next(items), target - *items.current, path1, result);
                search(ItemsRange::next(items), target, path, result);
            }
            return;
    }
}

int main(int argc, const char* argv[])
{
    Items_t items{ 2, 10, 13, 17, 22, 42 };
    Result_t result;
    int32_t target = 52;

    std::cout << "Input: "  << items << std::endl;
    std::cout << "Target: " << target << std::endl;
    search(ItemsRange{items}, target, Items_t{}, result);
    std::cout << "Output: " << result << std::endl;
    return 0;
}

【讨论】:

  • 是的,我应该要求退款这本书是垃圾。也感谢您的解释非常有帮助。
【解决方案3】:

代码正确地实现了算法,只是您没有在输出循环中应用从一开始的数组逻辑。变化:

for (int j = 0; j < arraySize; j++) {
    std::cout << include[j] << " ";
}

到:

for (int j = 0; j < arraySize; j++) {
    std::cout << include[j+1] << " ";
}

根据您组织代码的方式,确保在定义 sos 时定义 promising。

看到它在repl.it 上运行。输出:

0 1 0 0 0 1
0 0 1 1 1 0

该算法运行良好:sos 函数的第二个和第三个参数充当一个窗口,运行总和应保留在其中,promising 函数根据此窗口进行验证。此窗口之外的任何值要么太小(即使将所有剩余值都添加到其中,它仍然会小于目标值),或者太大(已经超出目标值)。这两个约束在本书第 5.4 章的开头进行了解释。

在每个索引处有两种可能的选择:包含总和中的值,或者不包含。 includes[i+1] 处的值表示此选择,并且都尝试过。当在这种递归尝试的深处存在匹配时,将输出所有这些选择(0 或 1)。否则,它们将被忽略并在第二次尝试中切换到相反的选择。

【讨论】:

    猜你喜欢
    • 2018-04-06
    • 1970-01-01
    • 2017-10-27
    • 1970-01-01
    • 1970-01-01
    • 2023-04-05
    • 2021-04-28
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多