【问题标题】:Generic Function remove() vs Member Function remove() for linked list链表的通用函数 remove() 与成员函数 remove()
【发布时间】:2018-10-31 12:58:05
【问题描述】:

我正在阅读 Nicolai M.Josuttis 撰写的“The C++ STL. A Tutorial and References”一书,在其中一章专门讨论 STL 算法的作者陈述如下: 如果你调用 remove() 列表的元素,算法不知道它正在对列表进行操作,因此执行它的操作 适用于任何容器:通过更改元素的值来重新排序元素。例如,如果算法 删除第一个元素,所有以下元素都分配给它们之前的元素。这 行为与列表的主要优点相矛盾:插入、移动和删除元素的能力 修改链接而不是值。为避免性能不佳,列表为所有操作算法提供了特殊的成员函数。你应该总是喜欢它们。此外,这些成员函数确实删除了“已删除”的元素,如下例所示:

#include <list>
#include <algorithm>
using namespace std;
int main()
{
list<int> coll;
// insert elements from 6 to 1 and 1 to 6
for (int i=1; i<=6; ++i) {
coll.push_front(i);
coll.push_back(i);
}
// remove all elements with value 3 (poor performance)
coll.erase (remove(coll.begin(),coll.end(),
3),
coll.end());
// remove all elements with value 4 (good performance)
coll.remove (4);
}

当然,这似乎足以令人信服,值得进一步考虑,但无论如何,我决定在我的 PC 中查看运行类似代码的结果,特别是在 MSVC 2013 环境中。 这是我的即兴代码:

int main()
{
    srand(time(nullptr));
    list<int>my_list1;
    list<int>my_list2;
    int x = 2000 * 2000;

    for (auto i = 0; i < x; ++i)
    {
        auto random = rand() % 10;
        my_list1.push_back(random);
        my_list1.push_front(random);
    }

    list<int>my_list2(my_list1);

    auto started1 = std::chrono::high_resolution_clock::now();
    my_list1.remove(5);
    auto done1 = std::chrono::high_resolution_clock::now();
    cout << "Execution time while using member function remove: " << chrono::duration_cast<chrono::milliseconds>(done1 - started1).count();

    cout << endl << endl;

    auto started2 = std::chrono::high_resolution_clock::now();
    my_list2.erase(remove(my_list2.begin(), my_list2.end(),5), my_list2.end());
    auto done2 = std::chrono::high_resolution_clock::now();
    cout << "Execution time while using generic algorithm remove: " << chrono::duration_cast<chrono::milliseconds>(done2 - started2).count();

    cout << endl << endl;
}

看到以下输出时我很惊讶:

Execution time while using member function remove: 10773

Execution time while using generic algorithm remove: 7459 

您能否解释一下这种矛盾行为的原因是什么?

【问题讨论】:

  • 在这种情况下,性能可能不是最合适的原因; a) 因为它可能是错误的,b) 因为 std::list 的性能是出了名的差(当然取决于用例)。听起来更相关的是,一种方法使指向您未删除的元素的指针和引用无效,而另一种方法则没有。毕竟,这样的保证通常是使用链表开始的原因。
  • int 非常小且易于复制。如果复制值比仅断开节点的成本更高,则这两种方法之间的差异可能会有所不同。
  • @Blastfurnace 这些值通常不会被复制,只是交换。这通常也很便宜。尤其是与列表遍历相比。
  • 您是在调试版本还是优化版本中测量性能?
  • @BaummitAugen 我认为std::remove 不会交换值(它们可能会被移动)。

标签: c++ c++11 linked-list stl erase-remove-idiom


【解决方案1】:

这是一个缓存问题。大多数性能问题都是缓存问题。我们总是想认为算法是首先要看的东西。但是,如果您故意强制编译器在一次运行中使用来自不同位置的内存,而在下一次运行时使用所有来自下一个位置的内存,您将看到缓存问题。

通过在构建原始列表时注释掉push_back 或push_front,我强制编译器创建代码以在my_list1 中构建具有连续内存元素的列表。

my_list2 始终位于连续内存中,因为它是在单个副本中分配的。

运行输出:

Execution time while using member function remove: 121

Execution time while using generic algorithm remove: 125


Process finished with exit code 0

这是我的代码,其中一个推送被注释掉了。

#include <list>
#include <algorithm>
#include <chrono>
#include <iostream>

using namespace std;

int main()
{
    srand(time(nullptr));
    list<int>my_list1;
//    list<int>my_list2;
    int x = 2000 * 2000;

    for (auto i = 0; i < x; ++i)
    {
        auto random = rand() % 10;
//        my_list1.push_back(random);  // avoid pushing to front and back to avoid cache misses.
        my_list1.push_front(random);
    }

    list<int>my_list2(my_list1);

    auto started1 = std::chrono::high_resolution_clock::now();
    my_list1.remove(5);
    auto done1 = std::chrono::high_resolution_clock::now();
    cout << "Execution time while using member function remove: " << chrono::duration_cast<chrono::milliseconds>(done1 - started1).count();

    cout << endl << endl;

    auto started2 = std::chrono::high_resolution_clock::now();
    my_list2.erase(remove(my_list2.begin(), my_list2.end(),5), my_list2.end());
    auto done2 = std::chrono::high_resolution_clock::now();
    cout << "Execution time while using generic algorithm erase: " << chrono::duration_cast<chrono::milliseconds>(done2 - started2).count();

    cout << endl << endl;
}

通过增加元素的数量并颠倒调用顺序以使擦除首先发生,然后删除发生,移除需要更长的时间。同样,这更多的是关于缓存而不是算法或正在完成的工作量。如果您运行的另一个程序会弄脏缓存、检查 Internet 或移动鼠标,那么您的 32 KB L1 缓存将被弄脏,并且该运行的性能会下降。

【讨论】:

  • 我只是将my_list1.push_back(random); 注释掉了。这可以防止代码从不同位置为列表的不同部分分配内存并提高缓存效率。
  • 随着x的增大,erase的效率提高,超过remove,但是缓存的效果还是很大的。
  • 通过将 x 乘以 8,然后将擦除的顺序切换为第一个,删除的顺序为第二个,我得到了 617 的擦除和 689 的删除的执行时间。记住,在原始测试中删除“更快”。但是,没有尝试运行删除秒。通过第二次运行删除并删除前后插入,删除速度较慢。
  • 如果你不相信这可能是一个缓存问题,那么你需要看这个:Scott Myers on Caching
猜你喜欢
  • 2017-12-18
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-03-21
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多