【问题标题】:Linux not respecting SCHED_FIFO priority ? ( normal or GDB execution )Linux 不尊重 SCHED_FIFO 优先级? (正常或 GDB 执行)
【发布时间】:2021-01-11 15:05:02
【问题描述】:

TL;DR

在多处理器/多核引擎上,可以在多个执行单元上调度多个 RT SCHED_FIFO 线程。因此优先级为 60 的线程和优先级为 40 的线程可以同时运行在 2 个不同的内核上。

这可能违反直觉,尤其是在模拟嵌入式系统时(通常像今天一样)在单核处理器上运行并依赖严格的优先级执行。

请参阅此帖子中的my other answer 了解摘要


原始问题描述

即使使用非常简单的代码来让 Linux 使用调度策略 SCHED_FIFO 尊重我的线程的优先级,我也遇到了困难。

  • 请参阅问题末尾的 MCVE。
  • 在答案中查看修改后的 MCVE

这种情况来自于需要在 Linux PC 下模拟嵌入式代码以执行集成测试

fifo 优先级10 的main 线程将启动线程divisor 和ratio。

divisor 线程应该得到priority 2,这样带有priority 1 的ratio 线程在b 得到一个合适的值之前不会评估a/b(这只是MCVE 的一个完全假设的场景,而不是真实的带有信号量或条件变量的生活案例)。

潜在的先决条件:您需要是 root 或更好地setcap 程序,以便可以更改调度策略和优先级

sudo setcap cap_sys_nice+ep main

johndoe@VirtualBox:~/Code/gdb_sched_fifo$ getcap main
main = cap_sys_nice+ep
  • 第一个实验是在 2 个 vCPU 的 Virtualbox 环境下进行的(gcc (Ubuntu 7.5.0-3ubuntu1~18.04) 7.5.0, GNU gdb (Ubuntu 8.1-0ubuntu3.2) 8.1.0.20180409-git),代码行为差不多OK 在正常执行下但NOK 在 GDB 下。

  • 本机 Ubuntu 20.04 上的其他实验显示了非常频繁的 NOK 行为,即使在 I3-1005 2C/4T (gcc (Ubuntu 9.3.0-10ubuntu2) 9.3.0、GNU gdb (Ubuntu 9.1-0ubuntu1) 9.1 的正常执行中) )

基本编译:

johndoe@VirtualBox:~/Code/gdb_sched_fifo$ g++ main.cc -o main -pthread

正常执行有时可以,有时如果没有 root 或没有 setcap 则不行

johndoe@VirtualBox:~/Code/gdb_sched_fifo$ ./main
Problem with setschedparam: Operation not permitted(1)  <<-- err msg if no root or setcap
Result: 0.333333 or Result: Inf                         <<-- 1/3 or div by 0

正常执行正常(例如使用 setcap )

johndoe@VirtualBox:~/Code/gdb_sched_fifo$ ./main
Result: 0.333333

现在如果你想调试这个程序,你会再次收到错误消息。

(gdb) run
Starting program: /home/johndoe/Code/gdb_sched_fifo/main 
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7f929a6a9700 (LWP 2633)]
Problem with setschedparam: Operation not permitted(1)     <<--- ERROR MSG
Result: inf                                                <<--- DIV BY 0
[New Thread 0x7f9299ea8700 (LWP 2634)]
[Thread 0x7f929a6a9700 (LWP 2633) exited]
[Thread 0x7f9299ea8700 (LWP 2634) exited]
[Inferior 1 (process 2629) exited normally]

这在这个问题gdb appears to ignore executable capabilities 中有解释(几乎所有答案都可能是相关的)。

所以在我的情况下,我做到了

  • sudo setcap cap_sys_nice+ep /usr/bin/gdb
  • 用set startup-with-shell off 创建一个~/.gdbinit

结果我得到了:

(gdb) run
Starting program: /home/johndoe/Code/gdb_sched_fifo/main 
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
[New Thread 0x7ffff6e85700 (LWP 2691)]
Result: inf                              <<-- NO ERR MSG but DIV BY 0 
[New Thread 0x7ffff6684700 (LWP 2692)]
[Thread 0x7ffff6e85700 (LWP 2691) exited]
[Thread 0x7ffff6684700 (LWP 2692) exited]
[Inferior 1 (process 2687) exited normally]
(gdb) 

结论和问题

  • 我认为唯一的问题来自 GDB
  • 在另一个(非虚拟)目标上的测试在正常执行下显示出更差的结果

我看到其他与 RT SCHED_FIFO 相关的问题没有得到尊重,但我发现答案没有或不明确的结论。我的 MCVE 也更小,潜在的副作用更少

Linux SCHED_FIFO not respecting thread priorities

SCHED_FIFO higher priority thread is getting preempted by the SCHED_FIFO lower priority thread?

评论带来了一些答案,但我仍然不相信......(......它应该像这样工作)

MCVE:

#include <iostream>
#include <thread>
#include <cstring>

double a = 1.0F;
double b = 0.0F;

void ratio(void)
{
    struct sched_param param;
    param.sched_priority = 1;
    int ret = pthread_setschedparam(pthread_self(),SCHED_FIFO,&param);
        if ( 0 != ret )
    std::cout << "Problem with setschedparam: " << std::strerror(errno) << '(' << errno << ')' << "\n" << std::flush;

    std::cout << "Result: " << a/b << "\n" << std::flush;
}

void divisor(void)
{
    struct sched_param param;
    param.sched_priority = 2;
    pthread_setschedparam(pthread_self(),SCHED_FIFO,&param);

    b = 3.0F;

    std::this_thread::sleep_for(std::chrono::milliseconds(2000u));
}


int main(int argc, char * argv[])
{
    struct sched_param param;
    param.sched_priority = 10;
    pthread_setschedparam(pthread_self(),SCHED_FIFO,&param);

    std::thread thr_ratio(ratio);
    std::thread thr_divisor(divisor);

    thr_ratio.join();
    thr_divisor.join();

    return 0;
}

【问题讨论】:

    标签: c++ linux c++11 gdb pthreads


    【解决方案1】:

    您的 MCVE 有几处明显错误:

    1. 您在 b 上存在数据竞争,即未定义的行为,因此任何事情都可能发生。

    2. 您期望divisor 线程将在pthread_setschedparam 调用之前完成ratio 线程开始计算比率。

      但绝对不能保证第一个线程不会在第二个线程创建之前很久就运行完成。

      这确实是在 GDB 下可能发生的事情:它必须捕获线程创建和销毁事件以跟踪所有线程,因此在 GDB 下创建线程明显慢于它之外.

    要解决第二个问题,请添加一个计数信号量,并让两个线程在 各自执行 pthread_setschedparam 调用之后进行循环。

    【讨论】:

    • 还有一些我不明白的地方。只有在比率线程优先级为 1 之后才应执行 a/b 操作。当时在系统上可能存在优先级为 10 的主线程(如果未在 join() 上阻塞)和优先级为 10 的除数线程(在 pthread_setchedparam(...) 调用之前)或 2(在 pthread_setchedparam(...) 调用之后)。所以我希望一旦比率线程的优先级为 1,它就会被阻塞,因为系统上有更高优先级的线程(和 SCHED_FIFO 策略)
    • @NGI 你用 1 个 CPU 配置你的 VBox 了吗?如果没有,主线程和第一个 (ratio) 线程将并行运行。
    • 好吧,我想我开始明白了……。我会看看。我认为有 2 个 vCPU。 (目前不在同一台计算机上)。反正我觉得很奇怪。我会接受 1 个 CPU 运行 sched_fifo RT 线程,1 个其他 CPU 运行 sched_rr RT 线程,其他 CPU 正常调度其他非 RT 线程。我会接受 2 个具有相同优先级的 sched_fifo RT 线程将在 2 个 CPU/核心上同时执行。但是我觉得有点奇怪的是,不同优先级的sched_fifo线程同时执行。这似乎是矛盾的。
    • 不同时执行它们将违反定义的行为。如果存在 SCHED_FIFO 1、SCHED_FIFO 2 和 SCHED_OTHER 线程,并且调度程序在两个 CPU 上运行 SCHED_FIFO 2 和 SCHED_OTHER,则违反了在 SCHED_OTHER 之前运行 SCHED_FIFO 1 的要求。
    • @TrentP。是的,对不起,你是对的。每个具有任何优先级 1..99 的 SCHED_FIFO RT 线程必须在任何其他具有策略 SCHED_OTHER 的线程可以运行之前完成或被阻塞。但这并没有改变我的审讯。当 SCHED_FIFO 2 已经在 CPU 上运行时,SCHED_FIFO 1 线程可以在另一个 CPU 上运行吗?如果我们小心为线程赋予更高的优先级,那是因为我们希望它在其他线程之前执行。我将在这个主题上扩展我对 SO 的搜索。
    【解决方案2】:

    我尝试了许多解决方案,但从未得到“无缺陷”代码。另请参阅此帖子中的my other answer

    最佳率,但不完美的代码是下面带有traditionnal pthread C的代码允许从一开始就创建具有正确属性的线程的语言。

    我仍然惊讶地发现即使使用此代码我仍然会出错(与 Question MCVE 相同,但使用纯 pthread... API)。

    为了强调代码我找到了以下序列

    $ seq 1000 | parallel ./main | grep inf
    Result: inf
    Result: inf
    ....
    

    inf 表示错误除以 0 结果。就我而言,缺陷约为 10/1000。

    像for i in {1..1000}; do ./main ; done | grep inf 这样的命令更长

    线程从高优先级到低优先级启动

    现在是除数线程

    • 先创建
    • 具有更高的 RT 优先级(2 > 1 > 主要停留在 SCHED_OTHER 非 RT 调度中)。

    所以我想知道为什么我仍然会被 0 除...

    最后我尝试减少任务集。运行正常

    $ taskset -pc 0 $$
    pid 2414's current affinity list: 0,1
    pid 2414's new affinity list: 0
    $ for i in {1..1000}; do ./main_oss ; done   <<-- no need for parallel in this case
    Result: 0.333333
    Result: 0.333333
    Result: 0.333333
    Result: 0.333333
    Result: 0.333333
    ...
    

    但是一旦 CPU 超过 1 个,缺陷就会再次出现

    $ taskset -pc 0,1 $$
    pid 2414's current affinity list: 0
    pid 2414's new affinity list: 0,1
    $ seq 1000 | parallel ./main_oss
    Result: 0.333333          | <<-- display by group of 2
    Result: 0.333333          |
    Result: inf             |   <<--
    Result: 0.333333        |
    ...
    

    当线程属于同一个父进程时,为什么我们在另一个CPU上运行优先级较低的RT SCHED_FIFO线程=?

    很遗憾,Linux 不支持 PTHREAD_SCOPE_PROCESS

    #include <iostream>
    #include <thread>
    #include <cstring>
    #include <pthread.h>
    
    double a = 1.0F;
    double b = 0.0F;
    
    void * ratio(void*)
    {
        std::cout << "Result: " << a/b << "\n" << std::flush;
        return nullptr;
    }
    
    void * divisor(void*)
    {
        b = 3.0F;
        std::this_thread::sleep_for(std::chrono::milliseconds(500u));
        return nullptr;
    }
    
    
    int main(int agrc, char * argv[])
    {
        struct sched_param param;
    
        pthread_t thr[2];
        pthread_attr_t attr;
        pthread_attr_init(&attr);
        pthread_attr_setschedpolicy(&attr,SCHED_FIFO);
        pthread_attr_setinheritsched(&attr,PTHREAD_EXPLICIT_SCHED);
    
        param.sched_priority = 2;
        pthread_attr_setschedparam(&attr,&param);
        pthread_create(&thr[0],&attr,divisor,nullptr);
    
        param.sched_priority = 1;
        pthread_attr_setschedparam(&attr,&param);
        pthread_create(&thr[1],&attr,ratio,nullptr);  
    
        pthread_join(thr[0],nullptr);
        pthread_join(thr[1],nullptr);
    
        return 0;
    } 
    

    【讨论】:

    • 您认为您的问题是否已解决?我打算看看它。我还想提一下,在 Linux 下,线程与进程 w.r.t 到内核没有区别,因此它可以很好地将它们调度到另一个 CPU 上
    • @Micrified,我的调试仍然有问题(使用 exec-wrapper 时设置断点的能力,使用 IDE 的能力......)。我将更新一个新的答案来收集这些数据)
    • 好的。好吧,我没有尝试过使用调试器,但是您实际上想要获得什么样的结果?我知道这是一个 MCVE,但您是否遇到了某种现实世界的应用程序/场景?
    • @Micrified。我添加了第二个"answer"。是的,当您想在 PC 上模拟单核嵌入式代码时,它具有实际用途(类似主题的其他问题总是提到这种情况)
    【解决方案3】:

    收集我在调试时遇到的剩余问题的新答案。

    Setting application affinity in gdb / Markus Ahlberg 之类的答案或gdb don't break when I use exec-wrapper script to exec my target binary 之类的问题提供了使用 GDB 选项 exec-wrapper 的解决方案,但后来我无法(总是)在我的代码中设置断点(甚至尝试我自己的包装器)

    我终于又回到了这个解决方案Setting application affinity in gdb / Craig Scratchley

    最初的问题

    $ ./main
    Result: inf
    

    运行时的解决方案

    taskset -c 0 ./main
    Result: 0.333333
    

    但用于调试

    gdb -ex 'set exec-wrapper taskset -c 0' ./main
    --> mixed result depending on conditions (native/virtualized ? Number of cores ? ) 
    sometimes 0.333333 sometimes inf
    --> problem to set breakpoints
    --> still work to do for me to summarize this issue
    

    或

    taskset -c 0 gdb main
    ...
    (gdb) r
    ...
    Result: inf
    

    最后

    taskset -c N chrt 99 gdb main <<-- where N is a core number (*)
    ...                           <<-- 99 denotes here "your higher prio in your system"
    (gdb) r
    ...
    Result: 0.333333
    
    • 我在上面写了 N 是因为如果你的程序 main 设置它与处理器 M 的亲和性,而你将 gdb 亲和性设置为 N,你可能会遇到同样的原始问题
    • 即使我对 SCHED_FIFO 而不是 SCHED_RR 感兴趣,我也只为 GDB 编写了 chrt 99,因为如果使用了选项 -f(用于 fifo),我经历了 gdb(或 IDE 见下文)冻结。我怀疑轮询机制更安全,因为线程总是会在某个时候释放

    如果你有一个 IDE(但不知道如何在这个 IDE 中正确设置 gdb)我可以做到

    taskset -c N chrt 99 code
    

    【讨论】:

      猜你喜欢
      • 2020-12-19
      • 2014-11-26
      • 2022-11-21
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多