【问题标题】:How to pause the timer when benchmarking an already multithreaded function in google benchmark?在谷歌基准测试中对已经多线程的函数进行基准测试时如何暂停计时器?
【发布时间】:2021-07-12 10:02:36
【问题描述】:

documentation on GitHub 有一个关于多线程基准测试的部分,但是,它需要将多线程代码放在基准测试定义中,并且库本身会使用多线程调用此代码。

我想对一个在内部创建线程的函数进行基准测试。我只对优化多线程部分感兴趣,所以我想单独对该部分进行基准测试。因此,我想在函数的顺序代码运行或内部线程正在创建/销毁时暂停计时器并进行设置/拆卸。

【问题讨论】:

    标签: multithreading google-benchmark


    【解决方案1】:

    使用线程屏障同步原语等到所有线程都已创建或完成设置等。此解决方案使用boost::barrier,但从C++20开始也可以使用std::barrier,或实现自定义屏障.自己实现时要小心,因为这很容易搞砸,但this answer 似乎是对的。

    将benchmark::State & state 传递给您的函数和线程,以便在需要时暂停/取消暂停。

    #include <thread>
    #include <vector>
    
    #include <benchmark/benchmark.h>
    #include <boost/thread/barrier.hpp>
    
    void work() {
        volatile int sum = 0;
        for (int i = 0; i < 100'000'000; i++) {
            sum += i;
        }
    }
    
    static void thread_routine(boost::barrier& barrier, benchmark::State& state, int thread_id) {
        // do setup here, if needed
        barrier.wait();  // wait until each thread is created
        if (thread_id == 0) {
            state.ResumeTiming();
        }
        barrier.wait();  // wait until the timer is started before doing the work
    
        // do some work
        work();
    
        barrier.wait();  // wait until each thread completes the work
        if (thread_id == 0) {
            state.PauseTiming();
        }
        barrier.wait();  // wait until the timer is stopped before destructing the thread
        // do teardown here, if needed
    }
    
    void f(benchmark::State& state) {
        const int num_threads = 1000;
        boost::barrier barrier(num_threads);
        std::vector<std::thread> threads;
        threads.reserve(num_threads);
        for (int i = 0; i < num_threads; i++) {
            threads.emplace_back(thread_routine, std::ref(barrier), std::ref(state), i);
        }
        for (std::thread& thread : threads) {
            thread.join();
        }
    }
    
    static void BM_AlreadyMultiThreaded(benchmark::State& state) {
        for (auto _ : state) {
            state.PauseTiming();
            f(state);
            state.ResumeTiming();
        }
    }
    
    BENCHMARK(BM_AlreadyMultiThreaded)->Iterations(10)->Unit(benchmark::kMillisecond)->MeasureProcessCPUTime(); // NOLINT(cert-err58-cpp)
    BENCHMARK_MAIN();
    

    在我的机器上,此代码输出(跳过标题):

    ---------------------------------------------------------------------------------------------
    Benchmark                                                   Time             CPU   Iterations
    ---------------------------------------------------------------------------------------------
    BM_AlreadyMultiThreaded/iterations:10/process_time       1604 ms       200309 ms           10
    

    如果我注释掉所有state.PauseTimer() / state.ResumeTimer(),它会输出:

    ---------------------------------------------------------------------------------------------
    Benchmark                                                   Time             CPU   Iterations
    ---------------------------------------------------------------------------------------------
    BM_AlreadyMultiThreaded/iterations:10/process_time       1680 ms       200102 ms           10
    

    我认为 80 毫秒的实时时间/200 毫秒的 CPU 时间差异在统计上是显着的,而不是噪声,这支持了这个示例正确工作的假设。

    【讨论】:

      猜你喜欢
      • 2017-05-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-02-02
      • 1970-01-01
      • 2023-03-13
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多