【问题标题】:C++ STL: Custom sorting one vector based on contents of another [duplicate]C ++ STL:根据另一个向量的内容自定义排序一个向量[重复]
【发布时间】:2010-12-15 22:27:24
【问题描述】:

这可能是最好的例子。我有两个向量/列表:

People = {Anne, Bob, Charlie, Douglas}
Ages   = {23, 28, 25, 21}

我想使用 sort(People.begin(), People.end(), CustomComparator) 之类的东西根据年龄对人物进行排序,但我不知道如何编写 CustomComparator 来查看年龄而不是人物。

【问题讨论】:

  • 这里的问题是 sort 会在People 中移动元素,但在Ages 中不会移动,因此会丢失存在的实际耦合。寻找建议将这两个属性组合在一起以避免丢失对应关系的各种答案。
  • 旁注:向量和列表是有区别的。您可以将 std::sort 用于向量,但应将特殊成员函数 std::list::sort 用于列表,因为它们没有随机访问迭代器。

标签: c++ sorting stl


【解决方案1】:

我建议将这两个列表合并为一个结构列表。这样你就可以像 Dirkgently 所说的那样简单地定义 operator <。

【讨论】:

    【解决方案2】:

    明显的方法

    通常的处理方法是创建一个包含姓名和年龄的对象的单个向量/列表,而不是创建两个单独的向量/列表:

    struct person { 
        std::string name;
        int age;
    };
    

    要根据年龄进行排序,请传递一个查看年龄的比较器:

    std::sort(people.begin(), people.end(), 
              [](auto const &a, auto const &b) { return a.age < b.age; });
    

    在较旧的 C++(C++11 之前,因此没有 lambda 表达式)中,您可以将比较定义为 operator&lt; 的成员重载或函数对象(重载 operator() 的对象)来执行比较:

    struct by_age { 
        bool operator()(person const &a, person const &b) const noexcept { 
            return a.age < b.age;
        }
    };
    

    那么你的排序看起来像:

    std::vector<person> people;
    // code to put data into people goes here.
    
    std::sort(people.begin(), people.end(), by_age());
    

    至于在为类定义operator&lt; 或使用上面显示的单独比较器对象之间进行选择,主要是一个问题,即是否存在对此类“显而易见”的单一排序。

    在我看来,按年龄对人进行分类并不一定很明显。但是,如果在您的程序上下文中很明显,除非您明确指定,否则按年龄对人进行排序,那么实现比较将是有意义的 作为person::operator&lt;,而不是像我上面那样在一个单独的比较类中。

    其他方法

    综上所述,在某些情况下,在排序之前将数据组合到结构中确实是不切实际或不可取的。

    如果是这种情况,您可以考虑几个选项。如果由于您使用的密钥太昂贵而无法交换(或者根本无法交换,尽管这种情况非常罕见),因此正常排序不切实际,那么您可能可以使用存储要排序的数据的类型以及与每个相关联的键集合的索引:

    using Person = std::pair<int, std::string>;
    
    std::vector<Person> people = {
        { "Anne", 0},
        { "Bob", 1},
        { "Charlie", 2},
        { "Douglas", 3}
    };
    
    std::vector<int> ages = {23, 28, 25, 21};
    
    std::sort(people.begin(), people.end(), 
        [](Person const &a, person const &b) { 
            return Ages[a.second] < Ages[b.second];
        });
    

    您还可以很容易地创建一个单独的索引,按键的顺序排序,然后使用该索引来读取关联的值:

    std::vector<std::string> people = { "Anne", "Bob", "Charlie", "Douglas" };   
    std::vector<int> ages = {23, 28, 25, 21};
    
    std::vector<std::size_t> index (people.size());
    std::iota(index.begin(), index.end(), 0);
    
    std::sort(index.begin(), index.end(), [&](size_t a, size_t b) { return ages[a] < ages[b]; });
    
    for (auto i : index) { 
        std::cout << people[i] << "\n";
    }
    

    但是请注意,在这种情况下,我们根本没有真正对项目本身进行排序。我们刚刚根据年龄对索引进行了排序,然后使用索引来索引我们想要排序的数据数组——但年龄和姓名都保持原来的顺序。

    当然,理论上你可能会遇到这样一种奇怪的情况,上面的方法都不起作用,你需要重新实现排序才能做你真正想做的事情。虽然我认为这种可能性可能存在,但我还没有在实践中看到它(我什至不记得看到我几乎认为这是正确的做法)。

    【讨论】:

    • 这一行应该是 std::sort(people.begin(), people.end(), by_age);我认为“by_age”之后不应该有括号
    • @jm1234567890:不是这样。 by_age 是一个类,我们需要的是一个对象。括号意味着我们正在调用它的 ctor,因此我们正在传递该类的一个对象来进行排序比较。但是,如果 by_age 是一个函数,那你就完全正确了。
    【解决方案3】:

    通常,您不会将想要保存在一起的数据放在不同的容器中。为 Person 创建一个结构/类并重载operator&lt;。

    struct Person
    {
        std::string name;
        int age;
    }
    
    bool operator< (const Person& a, const Person& b);
    

    或者如果这是一些扔掉的东西:

    typedef std::pair<int, std::string> Person;
    std::vector<Person> persons;
    std::sort(persons.begin(), persons.end());
    

    std::pair 已经实现了比较运算符。

    【讨论】:

    • 但是你会默认操作符
    • @Xavier:他让int 成为std::pair 的第一个成员。
    【解决方案4】:

    将它们保存在两个单独的数据结构中是没有意义的:如果您重新排序 People,您将不再有到 Ages 的合理映射。

    template<class A, class B, class CA = std::less<A>, class CB = std::less<B> >
    struct lessByPairSecond
        : std::binary_function<std::pair<A, B>, std::pair<A, B>, bool>
    {
        bool operator()(const std::pair<A, B> &left, const std::pair<A, B> &right) {
            if (CB()(left.second, right.second)) return true;
            if (CB()(right.second, left.second)) return false;
            return CA()(left.first, right.first);
        }
    };
    
    std::vector<std::pair<std::string, int> > peopleAndAges;
    peopleAndAges.push_back(std::pair<std::string, int>("Anne", 23));
    peopleAndAges.push_back(std::pair<std::string, int>("Bob", 23));
    peopleAndAges.push_back(std::pair<std::string, int>("Charlie", 23));
    peopleAndAges.push_back(std::pair<std::string, int>("Douglas", 23));
    std::sort(peopleAndAges.begin(), peopleAndAges.end(),
            lessByPairSecond<std::string, int>());
    

    【讨论】:

    • 这当然是有道理的,就像如果很少使用 Ages,那么每次引用 People 时,你都会为它的存在付出代价,因为缓存压力会增加
    【解决方案5】:

    正如其他人所指出的,您应该考虑将人员和年龄分组。

    如果您不能/不想这样做,您可以为它们创建一个“索引”,然后对该索引进行排序。例如:

    // Warning: Not tested
    struct CompareAge : std::binary_function<size_t, size_t, bool>
    {
        CompareAge(const std::vector<unsigned int>& Ages)
        : m_Ages(Ages)
        {}
    
        bool operator()(size_t Lhs, size_t Rhs)const
        {
            return m_Ages[Lhs] < m_Ages[Rhs];
        }
    
        const std::vector<unsigned int>& m_Ages;
    };
    
    std::vector<std::string> people = ...;
    std::vector<unsigned int> ages = ...;
    
    // Initialize a vector of indices
    assert(people.size() == ages.size());
    std::vector<size_t> pos(people.size());
    for (size_t i = 0; i != pos.size(); ++i){
        pos[i] = i;
    }
    
    
    // Sort the indices
    std::sort(pos.begin(), pos.end(), CompareAge(ages));
    

    现在,第n个人的名字是people[pos[n]],年龄是ages[pos[n]]

    【讨论】:

    • 非常感谢。请注意:这里的一些人将“人和年龄应该在同一个容器中”作为答案,而事实并非如此,并且有正当理由希望能够按照 OP 的要求去做。
    • +1。顺便说一下继承binary_function的目的是什么,是否可以去掉这个继承,仍然使用CompareAge?
    • @Kari: binary_function 自 C++11 起已弃用。在此之前,它的目的是提供一些函数适配器所需的 typedef。由于此处没有使用此类适配器,因此不需要,但通常的做法是定义这样的函子以确保它们与适配器兼容。
    【解决方案6】:

    Jerry Coffin 的回答非常清楚和正确。

    只是有一个相关的问题,可能会对该主题进行很好的讨论... :)

    我不得不根据向量的排序(比如序列)重新排序矩阵对象的列(比如TMatrix)... TMatrix 类不提供对其行的引用访问(因此我无法创建结构来对其重新排序......)但方便地提供了一个方法 TMatrix::交换(第 1 行,第 2 行)...

    这就是代码:

    TMatrix<double> matrix;
    vector<double> sequence;
    // 
    // 1st step: gets indexes of the matrix rows changes in order to sort by time
    //
    // note: sorter vector will have 'sorted vector elements' on 'first' and 
    // 'original indexes of vector elements' on 'second'...
    //
    const int n = int(sequence.size());
    std::vector<std::pair<T, int>> sorter(n);
    for(int i = 0; i < n; i++) {
        std::pair<T, int> ae;
        ae.first = sequence[i]; 
        ae.second = i;              
        sorter[i] = ae;
    }           
    std::sort(sorter.begin(), sorter.end());
    
    //
    // 2nd step: swap matrix rows based on sorter information
    //
    for(int i = 0; i < n; i++) {
        // updates the the time vector
        sequence[i] = sorter[i].first;
        // check if the any row should swap
        const int pivot = sorter[i].second;
        if (i != pivot) {
            //
            // store the required swaps on stack
            //
            stack<std::pair<int, int>> swaps;
            int source = pivot;
            int destination = i;
            while(destination != pivot) {
                // store required swaps until final destination 
                // is equals to first source (pivot)
                std::pair<int, int> ae;
                ae.first = source;
                ae.second = destination;
                swaps.push(ae);
                // retrieves the next requiret swap
                source = destination;
                for(int j = 0; j < n; j++) {
                    if (sorter[j].second == source) 
                        destination = j;
                        break;
                    }
                }
            }                   
            //
            // final step: execute required swaps
            //
            while(!swaps.empty()) {
                // pop the swap entry from the stack
                std::pair<int, int> swap = swaps.top();
                destination = swap.second;                      
                swaps.pop();
                // swap matrix coluns
                matrix.swap(swap.first, destination);
                // updates the sorter
                sorter[destination].second = destination;
            }
            // updates sorter on pivot
            sorter[pivot].second = pivot;
        }
    }
    

    我相信这仍然是 O(n log n),因为没有到位的每一行只会交换一次......

    玩得开心! :)

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2013-03-06
      • 1970-01-01
      • 1970-01-01
      • 2012-07-05
      • 1970-01-01
      • 2010-12-06
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多