【问题标题】:Reorder vector using a vector of indices [duplicate]使用索引向量重新排序向量[重复]
【发布时间】:2010-10-24 17:09:26
【问题描述】:

我想对向量中的项目重新排序,使用另一个向量来指定顺序:

char   A[]     = { 'a', 'b', 'c' };
size_t ORDER[] = { 1, 0, 2 };

vector<char>   vA(A, A + sizeof(A) / sizeof(*A));
vector<size_t> vOrder(ORDER, ORDER + sizeof(ORDER) / sizeof(*ORDER));

reorder_naive(vA, vOrder);
// A is now { 'b', 'a', 'c' }

以下是需要复制向量的低效实现:

void reorder_naive(vector<char>& vA, const vector<size_t>& vOrder)  
{   
    assert(vA.size() == vOrder.size());  
    vector vCopy = vA; // Can we avoid this?  
    for(int i = 0; i < vOrder.size(); ++i)  
        vA[i] = vCopy[ vOrder[i] ];  
}  

有没有更有效的方法,例如使用 swap()?

【问题讨论】:

  • 不要在函数名称中使用全部大写字母。或者任何事情,就此而言。只有#define 可以逃脱惩罚。
  • ORDER 的内容有歧义。是应该存储相应字母的索引,还是应该存储在该位置的字母的索引?对于给定的示例,两种解释都是正确的,尽管它们是不同的。
  • 我想到的是后者,但两种解释给出的结果完全相同。
  • 对于正在寻找上述reorder_naive 的有效版本的任何人,请勿使用下面提出的解决方案。他们计算问题的第一个,而不是后一个解释(参见上面的 cmets),但 DO NOT 提供相同的结果。

标签: c++ algorithm vector stl


【解决方案1】:

这个算法是基于 chmike 的,但是重排序索引的向量是const。此功能与他的所有 11 一致! [0..10] 的排列。复杂度为O(N^2),取N作为输入的大小,或者更准确地说,取最大orbit的大小。

关于修改输入的优化 O(N) 解决方案,请参见下文。

template< class T >
void reorder(vector<T> &v, vector<size_t> const &order )  {   
    for ( int s = 1, d; s < order.size(); ++ s ) {
        for ( d = order[s]; d < s; d = order[d] ) ;
        if ( d == s ) while ( d = order[d], d != s ) swap( v[s], v[d] );
    }
}

这是一个 STL 风格的版本,我投入了更多精力。它快了大约 47%(也就是说,几乎是 [0..10] 的两倍!),因为它尽可能早地完成所有交换,然后返回。重新排序向量由许多轨道组成,每个轨道在到达其第一个成员时重新排序。最后几个元素不包含轨道时会更快。

template< typename order_iterator, typename value_iterator >
void reorder( order_iterator order_begin, order_iterator order_end, value_iterator v )  {   
    typedef typename std::iterator_traits< value_iterator >::value_type value_t;
    typedef typename std::iterator_traits< order_iterator >::value_type index_t;
    typedef typename std::iterator_traits< order_iterator >::difference_type diff_t;
    
    diff_t remaining = order_end - 1 - order_begin;
    for ( index_t s = index_t(), d; remaining > 0; ++ s ) {
        for ( d = order_begin[s]; d > s; d = order_begin[d] ) ;
        if ( d == s ) {
            -- remaining;
            value_t temp = v[s];
            while ( d = order_begin[d], d != s ) {
                swap( temp, v[d] );
                -- remaining;
            }
            v[s] = temp;
        }
    }
}

最后,为了一劳永逸地回答这个问题,一个确实破坏了重新排序向量的变体(用-1填充它)。对于 [0..10] 的排列,它比之前的版本快 16% 左右。因为覆盖输入会启用动态规划,所以它是 O(N),对于某些具有较长序列的情况,渐近更快。

template< typename order_iterator, typename value_iterator >
void reorder_destructive( order_iterator order_begin, order_iterator order_end, value_iterator v )  {
    typedef typename std::iterator_traits< value_iterator >::value_type value_t;
    typedef typename std::iterator_traits< order_iterator >::value_type index_t;
    typedef typename std::iterator_traits< order_iterator >::difference_type diff_t;
    
    diff_t remaining = order_end - 1 - order_begin;
    for ( index_t s = index_t(); remaining > 0; ++ s ) {
        index_t d = order_begin[s];
        if ( d == (diff_t) -1 ) continue;
        -- remaining;
        value_t temp = v[s];
        for ( index_t d2; d != s; d = d2 ) {
            swap( temp, v[d] );
            swap( order_begin[d], d2 = (diff_t) -1 );
            -- remaining;
        }
        v[s] = temp;
    }
}

【讨论】:

  • @Potatoswatter:是的,这就是stackoverflow的工作方式......,当一个人进行(绝大多数)编辑时,从来没有真正理解过推理。就好像他们想阻止你改进你的答案或其他东西......
  • 您是否有任何示例代码表明此方法有效?我什至不能用 gcc 编译第二个版本。第一个版本没有正确重新排序我的向量。
  • @mangledorf: ideone.com/TsWbu - 请注意,重新排序向量包含数据向量中每个对应元素的最终位置,这与选择应该出现在哪个初始数据向量元素的重新排序向量不同每个最终位置。 OP 的例子在这方面是模棱两可的。我的代码可能可以调整为以其他方式工作,但我现在没有时间 :v( .
  • 即使元素已经在它的位置上,这个算法(至少是第二个版本)似乎也会交换。因此,如果您使用命令vector 调用它,即012,...它实际上会进行大量复制。
  • 我不打算测试这个,但我很确定非移动物体(奇异轨道)的优化是在第一个外循环的开头添加if ( s == order[ s ] ) continue;代码,或相同但在其他两个中用order_begin 替换order
【解决方案2】:

向量的就地重新排序

警告:排序索引的语义存在歧义。两者都在这里回答

将向量的元素移动到索引的位置

互动版here.

#include <iostream>
#include <vector>
#include <assert.h>

using namespace std;

void REORDER(vector<double>& vA, vector<size_t>& vOrder)  
{   
    assert(vA.size() == vOrder.size());

    // for all elements to put in place
    for( int i = 0; i < vA.size() - 1; ++i )
    { 
        // while the element i is not yet in place 
        while( i != vOrder[i] )
        {
            // swap it with the element at its final place
            int alt = vOrder[i];
            swap( vA[i], vA[alt] );
            swap( vOrder[i], vOrder[alt] );
        }
    }
}

int main()
{
    std::vector<double> vec {7, 5, 9, 6};
    std::vector<size_t> inds {1, 3,  0, 2};
    REORDER(vec, inds);
    for (size_t vv = 0; vv < vec.size(); ++vv)
    {
        std::cout << vec[vv] << std::endl;
    }
    return 0;
}

输出

9
7
6
5

请注意,您可以保存一个测试,因为如果有 n-1 个元素,最后第 n 个元素肯定就位。

退出时 vA 和 vOrder 已正确排序。

该算法最多执行 n-1 次交换,因为每次交换都会将元素移动到其最终位置。而且我们最多需要对 vOrder 进行 2N 次测试。

从索引的位置绘制vector的元素

以交互方式尝试here

#include <iostream>
#include <vector>
#include <assert.h>

template<typename T>
void reorder(std::vector<T>& vec, std::vector<size_t> vOrder)
{
    assert(vec.size() == vOrder.size());
            
    for( size_t vv = 0; vv < vec.size() - 1; ++vv )
    {
            if (vOrder[vv] == vv)
            {
                continue;
            }
            size_t oo;
            for(oo = vv + 1; oo < vOrder.size(); ++oo)
            {
                if (vOrder[oo] == vv)
                {
                    break;
                }
            }
            std::swap( vec[vv], vec[vOrder[vv]] );
            std::swap( vOrder[vv], vOrder[oo] );
    }
}

int main()
{
    std::vector<double> vec {7, 5, 9, 6};
    std::vector<size_t> inds {1, 3,  0, 2};
    reorder(vec, inds);
    for (size_t vv = 0; vv < vec.size(); ++vv)
    {
        std::cout << vec[vv] << std::endl;
    }
    return 0;
}

输出

5
6
7
9

【讨论】:

  • 你更新顺序向量的方法确实改进了代码,但是你真的需要内部的while循环吗?我相信您不会有以下理由:在算法的第一步中,您不需要 while 循环。它只需一步即可工作。有趣的是,您实际上是在减少原始问题。
  • @dribeas 不幸的是,这不是真的,但我一开始也是这么想的。尝试使用序列 2 3 1 0 你会发现 1 和 0 不会在正确的位置。
  • 它坏了,vA[i] 最终包含 vA[k] 其中k 是循环中的最后一个索引
  • 你能说得更清楚些吗?你指的是什么周期?我在代码中看到的唯一错误是 int 应该是 size_t。
  • 很抱歉更新旧线程,但这对我不起作用。我不得不使用 vOrder[i] 和 vOrder[vOrder[i]],而不是检查和使用 i 和 vOrder[i]。我在下面添加了一个答案。
【解决方案3】:

在我看来,vOrder 包含一组按所需顺序排列的索引(例如按索引排序的输出)。此处的代码示例遵循 vOrder 中的“循环”,其中跟随索引的子集(可能是所有 vOrder)将循环通过子集,并在子集的第一个索引处结束。

关于“循环”的维基文章

https://en.wikipedia.org/wiki/Cyclic_permutation

在以下示例中,每个交换都至少将一个元素放置在适当的位置。此代码示例根据 vOrder 有效地对 vA 重新排序,同时将 vOrder “无序”或“取消排列”返回到其原始状态 (0 :: n-1)。如果 vA 按顺序包含 0 到 n-1 的值,那么在重新排序后,vA 将在 vOrder 开始的地方结束。

template <class T>
void reorder(vector<T>& vA, vector<size_t>& vOrder)  
{   
    assert(vA.size() == vOrder.size());

    // for all elements to put in place
    for( size_t i = 0; i < vA.size(); ++i )
    { 
        // while vOrder[i] is not yet in place 
        // every swap places at least one element in it's proper place
        while(       vOrder[i] !=   vOrder[vOrder[i]] )
        {
            swap( vA[vOrder[i]], vA[vOrder[vOrder[i]]] );
            swap(    vOrder[i],     vOrder[vOrder[i]] );
        }
    }
}

这也可以使用移动而不是交换来更有效地实现。在移动过程中需要一个临时对象来保存一个元素。示例 C 代码,根据 I[] 中的索引对 A[] 重新排序,同时对 I[] 进行排序:

void reorder(int *A, int *I, int n)
{    
int i, j, k;
int tA;
    /* reorder A according to I */
    /* every move puts an element into place */
    /* time complexity is O(n) */
    for(i = 0; i < n; i++){
        if(i != I[i]){
            tA = A[i];
            j = i;
            while(i != (k = I[j])){
                A[j] = A[k];
                I[j] = j;
                j = k;
            }
            A[j] = tA;
            I[j] = j;
        }
    }
}

【讨论】:

  • sizeof(A) 是 sizeof(int*) 并且 sizeof(A[0]) 是 sizeof (int)
  • @QuentinUK - 答案现已修复。
【解决方案4】:

我认为,如果可以修改 ORDER 数组,那么对 ORDER 向量进行排序并在每次排序操作中交换相应值向量元素的实现就可以解决问题。

【讨论】:

    【解决方案5】:

    永远不要过早地优化。测量然后确定您需要优化的地方和内容。在许多性能不是问题的地方,您可以使用难以维护且容易出错的复杂代码结束。

    话虽如此,不要过早悲观。在不更改代码的情况下,您可以删除一半的副本:

        template <typename T>
        void reorder( std::vector<T> & data, std::vector<std::size_t> const & order )
        {
           std::vector<T> tmp;         // create an empty vector
           tmp.reserve( data.size() ); // ensure memory and avoid moves in the vector
           for ( std::size_t i = 0; i < order.size(); ++i ) {
              tmp.push_back( data[order[i]] );
           }
           data.swap( tmp );          // swap vector contents
        }
    

    此代码创建并清空(足够大)向量,其中按顺序执行单个副本。最后,交换有序向量和原始向量。这将减少副本,但仍需要额外的内存。

    如果您想就地执行移动,一个简单的算法可能是:

    template <typename T>
    void reorder( std::vector<T> & data, std::vector<std::size_t> const & order )
    {
       for ( std::size_t i = 0; i < order.size(); ++i ) {
          std::size_t original = order[i];
          while ( i < original )  {
             original = order[original];
          }
          std::swap( data[i], data[original] );
       }
    }
    

    应检查和调试此代码。简而言之,每个步骤中的算法将元素定位在第 i 个位置。首先,我们确定该位置的原始元素现在放置在数据向量中的位置。如果算法已经触及原始位置(它在第 i 个位置之前),则将原始元素交换到 order[original] 位置。再说一遍,那个元素可能已经被移动了......

    该算法在整数运算的数量上大约为 O(N^2),因此与初始 O(N) 算法相比,理论上性能时间更差。但是,如果 N^2 交换操作(最坏情况)的成本低于 N 复制操作,或者您确实受到内存占用的限制,它可以弥补。

    【讨论】:

    • 您的算法是最优的,因为 vOrder 不能被修改或者 N 很小。您可能会花更多时间扫描 vOrder,但只会交换 vA 值。我的交换 vOrder 和 vA 值最终可能比重新扫描具有小 N 值的 vOrder 更昂贵。
    • 嗯,不应该'tmp.resize(data.size());'在使用 'tmp.push_back(...)' 时是 'tmp.reserve(...)' 吗?要保持'resize',插入应该是'tmp[i] = ...'。
    【解决方案6】:

    现有答案调查

    你问是否有“更有效的方法”。但是您所说的高效是什么意思,您的要求是什么?

    Potatoswatter 的 answer 可以在 O(N²) 时间内使用 O(1) 额外空间,并且不会改变重新排序向量.

    chmikercgldr 给出的答案使用 O(N) 时间和 O(1) 额外空间,但他们通过改变重新排序向量来实现这一点.

    您的原始答案分配新空间,然后将数据复制到其中,而Tim MB 建议使用移动语义。然而,移动仍然需要一个地方来移动东西,并且像std::string 这样的对象既有长度变量也有指针。换句话说,基于移动的解决方案需要为任何对象分配O(N),为新向量本身分配O(1)。我在下面解释为什么这很重要。

    保留重新排序向量

    我们可能想要那个重新排序的向量!排序成本 O(N log N)。但是,如果您知道您将以相同的方式对多个向量进行排序,例如在 Structure of Arrays (SoA) 上下文中,您可以排序一次然后重复使用结果。这样可以节省很多时间。

    您可能还想对数据进行排序然后取消排序。拥有重新排序向量允许您执行此操作。这里的一个用例是在 GPU 上执行基因组测序,其中通过批量处理相似长度的序列来获得最大的速度效率。我们不能依赖用户按此顺序提供序列,因此我们先排序然后再取消排序。

    那么,如果我们想要所有世界中最好的怎么办:O(N) 处理没有额外分配的成本,但也没有改变我们的排序向量(我们毕竟,可能想要重用)?要找到那个世界,我们需要问:

    为什么多余的空间不好?

    您可能不想分配额外空间的原因有两个。

    首先是你没有太多的工作空间。这可能在两种情况下发生:您使用的是内存有限的嵌入式设备。通常这意味着您正在处理小型数据集,因此 O(N²) 解决方案在这里可能很好。但是,当您使用 非常 大型数据集时,也会发生这种情况。在这种情况下,O(N²) 是不可接受的,您必须使用一种 O(N) 变异解决方案。

    额外空间不好的另一个原因是分配很昂贵。对于较小的数据集,它可能比实际计算成本更高。因此,实现效率的一种方法是消除分配。

    大纲

    当我们改变排序向量时,我们这样做是为了表明元素是否在它们的置换位置。我们可以使用位向量来指示相同的信息,而不是这样做。但是,如果我们每次都分配位向量,那将是昂贵的。

    相反,我们可以通过将位向量每次重置为零来清除它。但是,这会导致每次函数使用的额外 O(N) 成本。

    相反,我们可以将“版本”值存储在向量中,并在每次使用函数时递增。这给了我们O(1) 访问权限,O(1) 清晰,以及摊销的分配成本。这类似于persistent data structure。不利的一面是,如果我们过于频繁地使用排序函数,则需要重置版本计数器,尽管这样做的 O(N) 成本已摊销。

    这就提出了一个问题:版本向量的最佳数据类型是什么?位向量最大化缓存利用率,但每次使用后都需要完全 O(N) 重置。 64 位数据类型可能永远不需要重置,但缓存利用率很低。实验是解决这个问题的最好方法。

    两种排列方式

    我们可以将排序向量视为具有两种意义:向前和向后。在前向意义上,向量告诉我们元素的去向。在向后的意义上,向量告诉我们元素来自哪里。由于排序向量隐含地是一个链表,向后的意义需要O(N) 额外的空间,但是,同样,我们可以摊销分配成本。依次应用这两种感官会让我们回到原来的顺序。

    性能

    在我的“Intel(R) Xeon(R) E-2176M CPU @ 2.70GHz”上运行单线程,对于长度为 32,767 个元素的序列,以下代码每次重新排序大约需要 0.81 毫秒。

    代码

    带有测试的两种感官的完整注释代码:

    #include <algorithm>
    #include <cassert>
    #include <random>
    #include <stack>
    #include <stdexcept>
    #include <vector>
    
    ///@brief Reorder a vector by moving its elements to indices indicted by another 
    ///       vector. Takes O(N) time and O(N) space. Allocations are amoritzed.
    ///
    ///@param[in,out] values   Vector to be reordered
    ///@param[in]     ordering A permutation of the vector
    ///@param[in,out] visited  A black-box vector to be reused between calls and
    ///                        shared with with `backward_reorder()`
    template<class ValueType, class OrderingType, class ProgressType>
    void forward_reorder(
      std::vector<ValueType>          &values,
      const std::vector<OrderingType> &ordering,
      std::vector<ProgressType>       &visited
    ){
      if(ordering.size()!=values.size()){
        throw std::runtime_error("ordering and values must be the same size!");
      }
    
      //Size the visited vector appropriately. Since vectors don't shrink, this will
      //shortly become large enough to handle most of the inputs. The vector is 1
      //larger than necessary because the first element is special.
      if(visited.empty() || visited.size()-1<values.size());
        visited.resize(values.size()+1);
    
      //If the visitation indicator becomes too large, we reset everything. This is
      //O(N) expensive, but unlikely to occur in most use cases if an appropriate
      //data type is chosen for the visited vector. For instance, an unsigned 32-bit
      //integer provides ~4B uses before it needs to be reset. We subtract one below
      //to avoid having to think too much about off-by-one errors. Note that
      //choosing the biggest data type possible is not necessarily a good idea!
      //Smaller data types will have better cache utilization.
      if(visited.at(0)==std::numeric_limits<ProgressType>::max()-1)
        std::fill(visited.begin(), visited.end(), 0);
    
      //We increment the stored visited indicator and make a note of the result. Any
      //value in the visited vector less than `visited_indicator` has not been
      //visited.
      const auto visited_indicator = ++visited.at(0);
    
      //For doing an early exit if we get everything in place
      auto remaining = values.size();
    
      //For all elements that need to be placed
      for(size_t s=0;s<ordering.size() && remaining>0;s++){
        assert(visited[s+1]<=visited_indicator);
    
        //Ignore already-visited elements
        if(visited[s+1]==visited_indicator)
          continue;
    
        //Don't rearrange if we don't have to
        if(s==visited[s])
          continue;
    
        //Follow this cycle, putting elements in their places until we get back
        //around. Use move semantics for speed.
        auto temp = std::move(values[s]);
        auto i = s;
        for(;s!=(size_t)ordering[i];i=ordering[i],--remaining){
          std::swap(temp, values[ordering[i]]);
          visited[i+1] = visited_indicator;
        }
        std::swap(temp, values[s]);
        visited[i+1] = visited_indicator;
      }
    }
    
    
    
    ///@brief Reorder a vector by moving its elements to indices indicted by another 
    ///       vector. Takes O(2N) time and O(2N) space. Allocations are amoritzed.
    ///
    ///@param[in,out] values   Vector to be reordered
    ///@param[in]     ordering A permutation of the vector
    ///@param[in,out] visited  A black-box vector to be reused between calls and
    ///                        shared with with `forward_reorder()`
    template<class ValueType, class OrderingType, class ProgressType>
    void backward_reorder(
      std::vector<ValueType>          &values,
      const std::vector<OrderingType> &ordering,
      std::vector<ProgressType>       &visited
    ){
      //The orderings form a linked list. We need O(N) memory to reverse a linked
      //list. We use `thread_local` so that the function is reentrant.
      thread_local std::stack<OrderingType> stack;
    
      if(ordering.size()!=values.size()){
        throw std::runtime_error("ordering and values must be the same size!");
      }
    
      //Size the visited vector appropriately. Since vectors don't shrink, this will
      //shortly become large enough to handle most of the inputs. The vector is 1
      //larger than necessary because the first element is special.
      if(visited.empty() || visited.size()-1<values.size());
        visited.resize(values.size()+1);
    
      //If the visitation indicator becomes too large, we reset everything. This is
      //O(N) expensive, but unlikely to occur in most use cases if an appropriate
      //data type is chosen for the visited vector. For instance, an unsigned 32-bit
      //integer provides ~4B uses before it needs to be reset. We subtract one below
      //to avoid having to think too much about off-by-one errors. Note that
      //choosing the biggest data type possible is not necessarily a good idea!
      //Smaller data types will have better cache utilization.
      if(visited.at(0)==std::numeric_limits<ProgressType>::max()-1)
        std::fill(visited.begin(), visited.end(), 0);
    
      //We increment the stored visited indicator and make a note of the result. Any
      //value in the visited vector less than `visited_indicator` has not been
      //visited.  
      const auto visited_indicator = ++visited.at(0);
    
      //For doing an early exit if we get everything in place
      auto remaining = values.size();
    
      //For all elements that need to be placed
      for(size_t s=0;s<ordering.size() && remaining>0;s++){
        assert(visited[s+1]<=visited_indicator);
    
        //Ignore already-visited elements
        if(visited[s+1]==visited_indicator)
          continue;
    
        //Don't rearrange if we don't have to
        if(s==visited[s])
          continue;
    
        //The orderings form a linked list. We need to follow that list to its end
        //in order to reverse it.
        stack.emplace(s);
        for(auto i=s;s!=(size_t)ordering[i];i=ordering[i]){
          stack.emplace(ordering[i]);
        }
    
        //Now we follow the linked list in reverse to its beginning, putting
        //elements in their places. Use move semantics for speed.
        auto temp = std::move(values[s]);
        while(!stack.empty()){
          std::swap(temp, values[stack.top()]);
          visited[stack.top()+1] = visited_indicator;
          stack.pop();
          --remaining;
        }
        visited[s+1] = visited_indicator;
      }
    }
    
    
    
    int main(){
      std::mt19937 gen;
      std::uniform_int_distribution<short> value_dist(0,std::numeric_limits<short>::max());
      std::uniform_int_distribution<short> len_dist  (0,std::numeric_limits<short>::max());
      std::vector<short> data;
      std::vector<short> ordering;
      std::vector<short> original;
    
      std::vector<size_t> progress;
    
      for(int i=0;i<1000;i++){
        const int len = len_dist(gen);
        data.clear();
        ordering.clear();
        for(int i=0;i<len;i++){
          data.push_back(value_dist(gen));
          ordering.push_back(i);
        }
    
        original = data;
    
        std::shuffle(ordering.begin(), ordering.end(), gen);
    
        forward_reorder(data, ordering, progress);
    
        assert(original!=data);
    
        backward_reorder(data, ordering, progress);
    
        assert(original==data);
      }  
    }
    

    【讨论】:

      【解决方案7】:

      使用 O(1) 空间要求进行重新排序是一项有趣的智力练习,但在 99.9% 的情况下,更简单的答案将满足您的需求:

      void permute(vector<T>& values, const vector<size_t>& indices)  
      {   
          vector<T> out;
          out.reserve(indices.size());
          for(size_t index: indices)
          {
              assert(0 <= index && index < values.size());
              out.push_back(std::move(values[index]));
          }
          values = std::move(out);
      }
      

      除了内存要求之外,我认为这会变慢的唯一方法是out 的内存与valuesindices 的内存位于不同的缓存页面中。

      【讨论】:

        【解决方案8】:

        你可以递归地做,我猜 - 像这样(未选中,但它给出了想法):

        // Recursive function
        template<typename T>
        void REORDER(int oldPosition, vector<T>& vA, 
                     const vector<int>& vecNewOrder, vector<bool>& vecVisited)
        {
            // Keep a record of the value currently in that position,
            // as well as the position we're moving it to.
            // But don't move it yet, or we'll overwrite whatever's at the next
            // position. Instead, we first move what's at the next position.
            // To guard against loops, we look at vecVisited, and set it to true
            // once we've visited a position.
            T oldVal = vA[oldPosition];
            int newPos = vecNewOrder[oldPosition];
            if (vecVisited[oldPosition])
            {
                // We've hit a loop. Set it and return.
                vA[newPosition] = oldVal;
                return;
            }
            // Guard against loops:
            vecVisited[oldPosition] = true;
        
            // Recursively re-order the next item in the sequence.
            REORDER(newPos, vA, vecNewOrder, vecVisited);
        
            // And, after we've set this new value, 
            vA[newPosition] = oldVal;
        }
        
        // The "main" function
        template<typename T>
        void REORDER(vector<T>& vA, const vector<int>& newOrder)
        {
            // Initialise vecVisited with false values
            vector<bool> vecVisited(vA.size(), false);
        
            for (int x = 0; x < vA.size(); x++)
            {
                REORDER(x, vA, newOrder, vecVisited);
            }
        }
        

        当然,你确实有 vecVisited 的开销。有人对这种方法有什么想法吗?

        【讨论】:

        • 在第一次阅读时,在我看来,这大致以相同的内存和时间成本结束。虽然您没有复制到另一个向量,但您复制到的局部变量与向量中的元素一样多。我必须更加努力地解释解释以测试正确性。
        • 是的,我也是这么想的。事实上,由于堆栈帧,它可能会占用更多空间,并且由于方法调用开销而需要更多时间。但无论如何,这是一个有趣的智力练习;-P
        【解决方案9】:

        遍历向量是 O(n) 操作。它有点难以击败。

        【讨论】:

        • 他说的是空间效率,而不是算法效率。
        【解决方案10】:

        您的代码已损坏。不能赋值给vA,需要使用模板参数。

        vector<char> REORDER(const vector<char>& vA, const vector<size_t>& vOrder)  
        {   
            assert(vA.size() == vOrder.size());  
            vector<char> vCopy(vA.size()); 
            for(int i = 0; i < vOrder.size(); ++i)  
                vCopy[i] = vA[ vOrder[i] ];  
            return vA;
        } 
        

        上面的效率稍微高一点。

        【讨论】:

        • OP 想要重新排序 vA 向量。所以原型应该是:void REORDER(vector& vA, const vector& vOrder).
        • 返回一个向量最终会导致必须创建一个副本,根据我的解释,这是 OP 想要避免的。您的方法的另一个优化不是创建大小元素的向量,而是保留内存并使用 push_backs。这将消除在复制之前默认构造所有元素的成本。
        【解决方案11】:

        标题和问题不清楚向量是否应该使用与订购 vOrder 相同的步骤进行订购,或者 vOrder 是否已经包含所需顺序的索引。 第一种解释已经有了令人满意的答案(参见 chmike 和 Potatoswatter),我对后者添加了一些想法。 如果对象 T 的创建和/或复制成本相关

        template <typename T>
        void reorder( std::vector<T> & data, std::vector<std::size_t> & order )
        {
         std::size_t i,j,k;
          for(i = 0; i < order.size() - 1; ++i) {
            j = order[i];
            if(j != i) {
              for(k = i + 1; order[k] != i; ++k);
              std::swap(order[i],order[k]);
              std::swap(data[i],data[j]);
            }
          }
        }
        

        如果您的对象的创建成本很小并且内存不是问题(请参阅 dribeas):

        template <typename T>
        void reorder( std::vector<T> & data, std::vector<std::size_t> const & order )
        {
         std::vector<T> tmp;         // create an empty vector
         tmp.reserve( data.size() ); // ensure memory and avoid moves in the vector
         for ( std::size_t i = 0; i < order.size(); ++i ) {
          tmp.push_back( data[order[i]] );
         }
         data.swap( tmp );          // swap vector contents
        }
        

        请注意,dribeas answer 中的两段代码做不同的事情。

        【讨论】:

          【解决方案12】:

          我试图使用@Potatoswatter 的解决方案将多个向量按第三个进行排序,并且对在犰狳sort_index 的索引输出向量上使用上述函数的输出感到非常困惑。要从sort_index(下面的arma_inds 向量)的向量输出切换到可以与@Potatoswatter 的解决方案(下面的new_inds)一起使用的向量输出,您可以执行以下操作:

          vector<int> new_inds(arma_inds.size());
          for (int i = 0; i < new_inds.size(); i++) new_inds[arma_inds[i]] = i;
          

          【讨论】:

            【解决方案13】:

            我想出了这个解决方案,它具有O(max_val - min_val + 1)空间复杂度,但它可以与std::sort 集成并受益于std::sortO(n log n) 不错时间复杂度

            std::vector<int32_t> dense_vec = {1, 2, 3};
            std::vector<int32_t> order = {1, 0, 2};
            
            int32_t max_val = *std::max_element(dense_vec.begin(), dense_vec.end());
            std::vector<int32_t> sparse_vec(max_val + 1);
            
            int32_t i = 0;
            for(int32_t j: dense_vec)
            {
                sparse_vec[j] = order[i];
                i++;
            }
            
            std::sort(dense_vec.begin(), dense_vec.end(),
                [&sparse_vec](int32_t i1, int32_t i2) {return sparse_vec[i1] < sparse_vec[i2];});
            

            在编写此代码时做出以下假设:

            • 向量值从零开始。
            • 向量不包含重复值。
            • 为了使用std::sort,我们有足够的内存可以牺牲

            【讨论】:

              【解决方案14】:

              这应该避免复制向量:

              void REORDER(vector<char>& vA, const vector<size_t>& vOrder)  
              {   
                  assert(vA.size() == vOrder.size()); 
                  for(int i = 0; i < vOrder.size(); ++i)
                      if (i < vOrder[i])
                          swap(vA[i], vA[vOrder[i]]);
              }
              

              【讨论】:

              • 我相信这已经被打破了。顺序:[ 2, 0, 1 ],第一步将交换第一个和第三个元素。第二步 i > order[i] 什么都不做,第三步 i > order[i] 再次什么都不做。
              • 此代码适用于示例,但不适用于 vOrder 序列 [3 0 1 2]。对于 i == 0 ,交换 3 和 2,则后续 i 值不会发生任何变化。
              • 感谢您指出错误。在发布代码之前,我应该检查更多。 :(
              猜你喜欢
              • 1970-01-01
              • 2014-11-13
              • 2018-06-21
              • 1970-01-01
              • 2015-09-01
              • 2015-12-31
              • 1970-01-01
              • 1970-01-01
              • 1970-01-01
              相关资源
              最近更新 更多