【问题标题】:How to deal with multiple threads in perl which turn into zombie如何处理perl中变成僵尸的多个线程
【发布时间】:2013-05-03 14:34:56
【问题描述】:

似乎在线程中使用管道可能会导致线程变成僵尸。事实上,管道中的命令变成了僵尸,而不是线程。这种情况不会发生,这很烦人,因为很难找出真正的问题。如何处理这个问题?这些是什么原因造成的?和管道有关吗?如何避免这种情况?

以下是创建示例文件的代码。

#buildTest.pl
use strict;
use warnings;

sub generateChrs{
    my ($outfile, $num, $range)=@_;
    open OUTPUT, "|gzip>$outfile";
    my @set=('A','T','C','G');
    my $cnt=0;
    while ($cnt<$num) {
        # body...
        my $pos=int(rand($range));
        my $str = join '' => map $set[rand @set], 1 .. rand(200)+1;
        print OUTPUT "$cnt\t$pos\t$str\n";
        $cnt++
    }
    close OUTPUT;
}

sub new_chr{
    my @chrs=1..22;
    push @chrs,("X","Y","M", "Other");
    return @chrs;
}

for my $chr (&new_chr){
    generateChrs("$chr.gz",50000,100000)
}

以下代码偶尔会创建僵尸线程。原因或触发因素仍然未知。

#paralRM.pl
use strict;
use threads;
use Thread::Semaphore;
my $s = Thread::Semaphore->new(10);

sub rmDup{
    my $reads_chr=$_[0];
    print "remove duplication $reads_chr START TIME: ",`date`;
    return 0 if(!-s $reads_chr);

    my $dup_removed_file=$reads_chr . ".rm.gz";
    $s->down();
    open READCHR, "gunzip -c $reads_chr |sort -n -k2 |" or die "Error: cannot open $reads_chr";
    open OUTPUT, "|sort -k4 -n|gzip>$dup_removed_file";

    my ($last_id, $last_pos, $last_reads)=split('\t',<READCHR>);
    chomp($last_reads);
    my $last_length=length($last_reads);
    my $removalCnts=0;

    while (<READCHR>) {
        chomp;
        my @line=split('\t',$_);
        my ($id, $pos, $reads)=@line;
        my $cur_length=length($reads);
        if($last_pos==$pos){
            #may dup
            if($cur_length>$last_length){
                ($last_id, $last_pos, $last_reads)=@line;
                $last_length=$cur_length;
            }
            $removalCnts++;
            next;
        }else{
            #not dup
        }
        print OUTPUT join("\t",$last_id, $last_pos, $last_reads, $last_length, "\n");
        ($last_id, $last_pos, $last_reads)=@line;
        $last_length=$cur_length;
    }

    print OUTPUT join("\t",$last_id, $last_pos, $last_reads, $last_length, "\n");
    close OUTPUT;
    close READCHR;
    $s->up();
    print "remove duplication $reads_chr END TIME: ",`date`;
    #unlink("$reads_chr")
    return $removalCnts;
}


sub parallelRMdup{
    my @chrs=@_;
    my %jobs;
    my @removedCnts;
    my @processing;

    foreach my $chr(@chrs){
        while (${$s}<=0) {
            # body...
            sleep 10;
        }
        $jobs{$chr}=async {
            return &rmDup("$chr.gz")
            }
        push @processing, $chr;
    };

    #wait for all threads finish
    foreach my $chr(@processing){
        push @removedCnts, $jobs{$chr}->join();
    }
}

sub new_chr{
    my @chrs=1..22;
    push @chrs,("X","Y","M", "Other");
    return @chrs;
}

&parallelRMdup(&new_chr);

【问题讨论】:

  • 您的所有线程是否都报告了合理的开始和结束时间?但我看不出您的代码有任何明显错误可能导致线程未加入。但是,也有一些不好的做法:①你是不是在async 块后面少了一个分号? ② 产生线程时不要忙等待。并且不要取消引用 Semaphore 对象。相反,您可以在生成之前 down 信号量,但 up 在线程末尾→ 更好。 ③ 你应该以编程方式断言所有@chrs 都是唯一的,否则你只会加入$chr 的最后一个线程。
  • 僵尸是在管道中创建的(排序、gzip 等)。谢谢你的建议。我学到了很多!

标签: multithreading perl


【解决方案1】:

正如您原始帖子中的 cmets 所建议的那样 - 您的代码没有任何明显错误。可能有助于理解zombie 进程是什么。

具体来说 - 它是一个衍生进程(由您的open)已退出,但父进程尚未收集它的返回码。

对于短时间运行的代码,这并不是那么重要 - 当您的主程序退出时,僵尸将“重新成为”init,这将自动清理它们。

为了更长时间的运行,您可以使用waitpid 来清理它们并收集返回码。

现在在这种特定情况下 - 我看不出具体问题,但我猜这与您打开文件句柄的方式有关。像你一样打开文件句柄的缺点是它们是全局范围的——当你做一些棘手的事情时,这通常是个坏消息。

我想如果您将 open 调用更改为:

my $pid = open ( my $exec_fh, "|-", "executable" ); 

然后在你的close 之后在$pid 上调用waitpid,然后你的僵尸就会完成。测试来自waitpid 的返回值,以了解您的哪些高管犯了错误(如果有的话),这应该可以帮助您找出原因。

或者 - 设置 $SIG{CHLD} = "IGNORE"; 这意味着你 - 有效地 - 告诉你的子进程“立即离开” - 但如果它们死了,你将无法从它们那里获得返回码。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-09-19
    • 2016-07-03
    • 2022-01-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多