【问题标题】:the relationship between thread running time, cpu context switching and performance线程运行时间、cpu上下文切换和性能的关系
【发布时间】:2018-06-12 22:38:42
【问题描述】:

我做了一个实验来模拟我们的服务器代码中发生的事情,我启动了 1024 个线程,每个线程都执行一个系统调用,这大约需要 2.8 秒才能在我的机器上完成执行。然后我在每个线程的函数中添加usleep(1000000),执行时间增加到16s,当我第二次运行相同的程序时时间将减少到8s。我猜这可能是由 cpu 缓存和 cpu 上下文切换引起的,但我不太确定如何解释。

此外,避免这种情况发生的最佳做法是什么(稍微增加每个线程的运行时间会导致整个程序性能下降)。

我在这里附上测试代码,非常感谢您的帮助。

//largetest.cc
#include "local.h"
#include <time.h>
#include <thread>
#include <string>
#include "unistd.h"

using namespace std;

#define BILLION 1000000000L

int main()
{

    struct timespec start, end;
    double diff;

    clock_gettime(CLOCK_REALTIME, &start);

    int i = 0;
    int reqNum = 1024;

    for (i = 0; i < reqNum; i++)
    {
        string command = string("echo abc");
        thread{localTaskStart, command}.detach();
    }

    while (1)
    {
        if ((localFinishNum) == reqNum)
        {
            break;
        }
        else
        {
            usleep(1000000);
        }
        printf("curr num %d\n", localFinishNum);
    }

    clock_gettime(CLOCK_REALTIME, &end); /* mark the end time */
    diff = (end.tv_sec - start.tv_sec) * 1.0 + (end.tv_nsec - start.tv_nsec) * 1.0 / BILLION;
    printf("debug for running time = (%lf) second\n", diff);

    return 0;
}
//local.cc
#include "time.h"
#include "stdlib.h"
#include "stdio.h"
#include "local.h"
#include "unistd.h"
#include <string>
#include <mutex>

using namespace std;

mutex testNotifiedNumMtx;
int localFinishNum = 0;

int localTaskStart(string batchPath)
{

    char command[200];

    sprintf(command, "%s", batchPath.data());

    usleep(1000000);

    system(command);

    testNotifiedNumMtx.lock();
    localFinishNum++;
    testNotifiedNumMtx.unlock();

    return 0;
}

//local.h


#ifndef local_h
#define local_h

#include <string>

using namespace std;

int localTaskStart( string batchPath);

extern int localFinishNum;
#endif

【问题讨论】:

    标签: c++ linux multithreading performance context-switching


    【解决方案1】:

    localFinishNum 的读取也应受 mutex 保护,否则根据线程被调度的位置(即在哪些内核上)、缓存何时以及如何失效等因素,结果是不可预测的.

    事实上,如果编译器决定将localFinishNum 放入寄存器(而不是总是从内存中加载),那么在优化模式下编译程序甚至可能不会终止。

    【讨论】:

      猜你喜欢
      • 2014-02-20
      • 1970-01-01
      • 2023-02-13
      • 2016-07-04
      • 2017-09-09
      • 1970-01-01
      • 2011-07-23
      • 1970-01-01
      相关资源
      最近更新 更多