【问题标题】:How does memory usage of thread_local scale with number of threads?thread_local 的内存使用量如何随线程数增加?
【发布时间】:2021-08-08 02:36:55
【问题描述】:

我认为 C/C++ 标准没有说明复杂性,所以我对具体的实现很好奇(我认为它们都有相同的行为)。

假设我有以下 C++ 函数。

void fn() {
    thread_local char arr[1024*1024]{};
    // do something with arr
}

我的程序有 80 个线程,其中 47 个线程至少运行一次 fn()。

我的程序的内存使用量是否增长了大约 47 倍某个常数,80 倍某个常数,还是有其他公式可以解决这个问题?

注意:this Java 问题由于某种原因已关闭,但如果 Java 使用与 C/C++ 相同的原语,则 IDK。

【问题讨论】:

  • 这是实现定义的,也取决于应用程序。在使用请求分页和过度使用的操作系统上,操作系统不会尝试为数组查找内存,直到尝试访问它时出现页面错误。但是,在其他操作系统上,当线程启动时,每个新创建的线程都必须为其所有本地内存小马。
  • 如果你认为你需要线程本地——你不需要。
  • 您有具体的实现方案吗? gcc、clang、msvc 等?
  • 'Thread local' 是线程的全局变量:(
  • @AyxanHaqverdili 我假设它们都使用相同的底层机制,但如果不是 clang,gcc msvc,在 Win 和 Linux 上。

标签: c++ c thread-local-storage


【解决方案1】:

根据C++11标准:

3.7.2 线程存储时长 [ basic.stc.thread ]

1 所有使用 thread_local 关键字声明的变量都有线程存储持续时间。这些存储 实体应在创建它们的线程期间持续存在。有一个明确的对象或 每个线程的引用,并且使用声明的名称是指与当前线程关联的实体。

2 具有线程存储持续时间的变量应在其第一次使用 odr (3.2) 之前进行初始化,如果已构建, 应在线程退出时销毁。

它说,“这些实体的存储将持续到创建它们的线程的持续时间。”。 所以,就我的阅读而言,内存必须分配给所有的线程。

但是,如果使用它们,它们只会被初始化和破坏:“具有线程存储持续时间的变量应在其第一次使用odr之前被初始化( 3.2) 并且,如果构造,应在线程退出时销毁”。

【讨论】:

    【解决方案2】:

    这可能很大程度上取决于实现,尽管您可以相当轻松地验证实现的行为。例如在 Windows 上运行以下程序(使用调试 Visual Studio 构建以避免优化删除未使用的代码):

    #include <iostream>
    #include <array>
    #include <thread>
    
    struct Foo
    {
        std::array<char, 1'000'000'000> data;
    };
    
    void bar()
    {
        thread_local Foo foo;
        for (int i = 0; i < foo.data.size(); i++)
        {
            foo.data[i] = i;
        }
        std::this_thread::sleep_for(std::chrono::seconds(1000));
    }
    
    int main()
    {
        std::thread thread1([]
        {
            bar();
        });
    
        std::thread thread2([]
        {
            std::this_thread::sleep_for(std::chrono::seconds(1000));
        });
    
        thread1.join();
        thread2.join();
    }
    

    使用 3GB 内存(1GB 用于两个线程,1GB 用于主线程)。删除 thread2 会将内存使用量降至 2GB。在 Linux 上,这种行为可能会有所不同,因为它存在过度分配,并且未使用的内存页面在使用之前不会被分配。

    您可以通过使用智能指针仅在实际使用时分配内存来避免这种情况,例如将bar 更改为:

    void bar()
    {
        thread_local std::unique_ptr<Foo> foo = std::make_unique<Foo>();
        for (int i = 0; i < foo->data.size(); i++)
        {
            foo->data[i] = i;
        }
        std::this_thread::sleep_for(std::chrono::seconds(1000));
    }
    

    将内存使用量减少到 1GB,因为只有 thread1 实际分配了大数组 thread2,而主线程只需要存储 unique_ptr。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-04-07
      • 2020-11-03
      • 1970-01-01
      • 1970-01-01
      • 2021-05-02
      • 2020-02-04
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多