【问题标题】:Perl substitution not workingPerl 替换不起作用
【发布时间】:2015-02-21 19:23:49
【问题描述】:

我有以下代码:

use warnings;
use strict;

my %hash;

open TRADUCTOR, $ARGV[0];
open GTF, $ARGV[1];

while (my $line = <TRADUCTOR>) {
    chomp $line;
    my @chompline = split(/\t/, $line);
    $hash{$chompline[2]} = $chompline[1];
}

while (my $line = <GTF>) {
    my @chompline1 = split(/\t/, $line);
    my @chompline2 = split(/ /, $chompline1[8]);
    my @chompline3 = split(/;/, $chompline2[2]);
    my $transcript = $chompline3[0];
    $transcript =~ s/"//g;
    $transcript =~ s/PAC://g;
    my $transcript2 = $transcript;
    $transcript2 =~ s/(_[jox])/&&&$1/g;
    my @chompline4 = split(/&&&/, $transcript2);
    my $coletilla = $chompline4[1];
    my $transl = $hash{$chompline4[0]};
    if (defined $chompline4[1]) {
        $line =~ s/PAC:$transcript;/$transl.$coletilla;/ee;
    } else {
        $line =~ s/PAC:$transcript;/$hash{$chompline4[0]};/ee;
    }
    print $line;
}

第一个参数是以下文件(前十行):

Sb01g017490 Sb01g017490.1   1951419
Sb02g039360 Sb02g039360.1   1959410
Sb01g037620 Sb01g037620.1   1953645
Sb03g003880 Sb03g003880.1   1960464
Sb01g001330 Sb01g001330.1   1949441
Sb01g049890 Sb01g049890.1   1955138
Sb09g030646 Sb09g030646.1   1982110
Sb02g011950 Sb02g011950.1   1956744
Sb04g008540 Sb04g008540.1   1965938

第二个参数是以下文件(前十行):

A01 greenc1.0   exon    5409    5518    .   -   .   gene_id "Bra.XLOC_002074";transcript_id "Bra.TCONS_00002741";transcript_biotype "protein_coding"
A01 greenc1.0   exon    5616    5654    .   -   .   gene_id "Bra.XLOC_002074";transcript_id "Bra.TCONS_00002741";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8307    8530    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627_j.1";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8426    8530    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8599    8844    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627_j.1";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8599    8823    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8919    9056    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627_j.1";transcript_biotype "protein_coding"
A01 greenc1.0   exon    8919    9056    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627";transcript_biotype "protein_coding"
A01 greenc1.0   exon    9151    9413    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627_j.1";transcript_biotype "protein_coding"
A01 greenc1.0   exon    9151    9413    .   -   .   gene_id "Bra011902";transcript_id "PAC:22703627";transcript_biotype "protein_coding"

我想将文件 2 中的字符串替换为来自文件 1 的另一个字符串。但是,文本替换不起作用...没有任何内容被替换。怎么回事?

这是有问题的代码:

if (defined $chompline4[1]) {
    $line =~ s/PAC:$transcript;/$transl.$coletilla;/ee;
} else {
    $line =~ s/PAC:$transcript;/$hash{$chompline4[0]};/ee;
}

【问题讨论】:

  • Sb01g017490 Sb01g017490.1 1951419 中有实际的标签吗?
  • 关于/ee双重评估的使用,看这个问题有一个可怕的例子:stackoverflow.com/q/9591658/725418

标签: regex string perl replace


【解决方案1】:

您在替换正则表达式中使用/ee。这是一种可怕的做法,只能作为 Perl 被滥用的令人沮丧的例子。或者鼓励有经验的(?)程序员如何滥用 Perl 的例子。

您不需要使用评估来插入变量。您不需要评估来连接字符串。而您正在评估 两次,这意味着:

$foo = "bar";   # assume
s/.../$foo/ee   # before
s/.../bar/e     # first eval
s/.../bar/      # second eval - bareword becomes string, if you are lucky

第二次评估时发出的警告与您在程序中输入bar 相同:

Unquoted string "bar" may clash with future reserved word
Name "main::bar" used only once: possible typo

除非您希望自己的数据被破坏,否则我建议您使用不同的策略。您当前的代码允许部分匹配,删除分隔符,并且在应该考虑的地方没有考虑引号。我认为这就是为什么您要修改数据以准备此正则表达式替换的原因。

我会给你一些关于如何解决问题的建议,但我觉得我没有足够的有效信息来说明你正在尝试做什么,除了看起来你正在尝试替换例如

"Bra011902";transcript_id "PAC:22703627";transcript_biotype "protein_coding"

与

Sb01g017490.1   

【讨论】:

    【解决方案2】:

    我试过运行你的程序,但是很困难,因为我不知道哪些空格是空格,哪些是制表符。

    我建议不要使用诸如chomplines 之类的随机名称,而是使用实际的字段名称。此外,使用/\s*/ 打破您的字段,这将允许您同时获取制表符和空格:

    chomp $line; # You forgot it the first time around!
    my ($f1, $f2, $f3, $f4, $f5, $f6...) = split /\s*/, $line;
    

    其中$f1、$f2 等是要拆分的字段的实际名称。

    我个人会创建一个由字段值作为键的哈希:

    my @fields = qw(field1 field2 field3 field4 field5 field6); 
    

    然后,拆分该行并从中创建一个散列。 (我很想用map 想办法做到这一点):

    my %values;
    my @field_values = split /\s+/, $line;
    for my $index ( (0..$#fields) ) {
        $values{ $field[$index] } = $field_values[$index];
    }
    

    这将使您的程序更容易调试,并且更容易跟踪正在发生的事情。另外,通过使用/\s+/ 在空白 上拆分一次,您不必担心制表符和空格会被弄乱,或者两个制表符分隔单个字段。您将需要第二次拆分来拆分用分号分隔的字段,但使用字段名称而不是数组位置会使事情更容易处理。

    【讨论】:

      猜你喜欢
      • 2014-05-24
      • 1970-01-01
      • 2012-01-18
      • 1970-01-01
      • 1970-01-01
      • 2018-05-02
      • 2018-04-27
      • 2018-03-15
      • 2013-03-09
      相关资源
      最近更新 更多