【问题标题】:Check for decimal point and add it at the end if its not there using awk/perl使用 awk/perl 检查小数点,如果不存在则在末尾添加
【发布时间】:2015-09-01 06:28:38
【问题描述】:

我有 test.dat 文件,其值如下:

    20150202,abc,,,,3625.300000,,,,,-5,,,,,,,,,,,,,,,,,,,,,,
    20150202,def,,,,32.585,,,,,0,,,,,,,,,,,,,,,,,,,,,,
    20150202,xyz,,,,12,,,,,0.004167,,,,,,,,,,,,,,,,,,,,,,

我的预期输出如下所示:

   20150202,abc,,,,3625.300000,,,,,-5.,,,,,,,,,,,,,,,,,,,,,,
                                     ^. added here
   20150202,def,,,,32.585,,,,,0.,,,,,,,,,,,,,,,,,,,,,,
                               ^. added here
   20150202,xyz,,,,12.,,,,,0.004167,,,,,,,,,,,,,,,,,,,,,,
                     ^. added here

所以如果第 6 列和第 11 列没有小数点,那么我们应该添加 '.'在文件末尾。

我尝试了下面的代码,但它在拆分期间抛出错误消息

    #!/usr/bin/perl
    use strict;
    use warnings;
    my $filename = 'test.dat';
    open my $fh, $filename or die "Could not open file '$filename': $!";
    my @cols_to_change = qw ( 6 11 );
    while (my $val = <$fh>) {
       my @row = split (/,/);
       foreach my $col ( @cols_to_change ) {
          unless ( $row[$col] =~ m/\./ ) { $row[$col] .= '.' }
       }
    print join ( ',', @row );
    }

我收到的错误信息如下:

Use of uninitialized value in split at test.pl line 11, <$fh> line 1.
Use of uninitialized value in pattern match (m//) at test.pl line 13, <$fh> line 1.
Use of uninitialized value in pattern match (m//) at test.pl line 13, <$fh> line 1.
Use of uninitialized value in join or string at test.pl line 15, <$fh> line 1.
....

我不允许使用任何额外的 perl 模块,例如 Text::CSV。此外,任何使用 awk 的解决方案都会有很大帮助!

【问题讨论】:

标签: regex perl shell awk


【解决方案1】:

这是你的错误:

while (my $val = <$fh>) {

将其更改为:

while ( <$fh> ) {

或告诉split 使用$val。

split ( /,/, $val ); 

在上面的代码中 - 您在 while 循环中设置了 $val,但各种模式匹配和拆分根本没有使用 $val。

另请注意 - perl 数组中的第一个元素是 0。因此,您可能应该保留您复制的示例代码中的 5 和 10。

https://stackoverflow.com/questions/30847880/how-to-check-whether-number-has-decimal-point-in-it-and-add-decimal-point-at-the/30848190#30848190

这个:

#!/usr/bin/perl
use strict;
use warnings;
my $filename = 'test.dat';
open my $fh, $filename or die "Could not open file '$filename': $!";
my @cols_to_change = qw ( 5 10 );
while (<$fh>) {
    my @row = split(/,/);
    foreach my $col (@cols_to_change) {
        unless ( $row[$col] =~ m/\./ ) { $row[$col] .= '.' }
    }
    print join( ',', @row );
}

输出:

20150202,abc,,,,3625.300000,,,,,-5.,,,,,,,,,,,,,,,,,,,,,,
20150202,def,,,,32.585,,,,,0.,,,,,,,,,,,,,,,,,,,,,,
20150202,xyz,,,,12.,,,,,0.004167

根据要求,在输入您的样本数据时。

【讨论】:

    【解决方案2】:

    在 awk 中

    只为这些字段添加子

    awk -F, -vOFS="," '{sub(/^[^\.]+$/,"&.",$6);sub(/^[^\.]+$/,"&.",$11)}1' file
    

    或sed

    sed 's/^\(\([^,]*,\)\{5\}[^.,]\+\),/\1./;s/^\(\([^,]*,\)\{10\}[^.,]\+\),/\1./' file
    

    【讨论】:

      【解决方案3】:

      为了完整性,

      1. 狂欢

        (
            IFS=,
            while read -ra f; do 
                for i in 5 10; do 
                    [[ ${f[i]} == *.* ]] || f[i]+=.   # add a dot if not there
                done
                echo "${f[*]}"                        # quotes required here
            done < file
        )
        

        括号将代码放在一个子shell中,因此 IFS 的值不会在您当前的shell中更改。

      2. sed

        sed -r '
            s/^(([^,]*,){5})([^.,]+),/\1\3.,/
            s/^(([^,]*,){10})([^.,]+),/\1\3.,/
        ' file
        

      【讨论】:

        猜你喜欢
        • 2018-07-05
        • 2022-01-02
        • 1970-01-01
        • 2022-12-11
        • 1970-01-01
        • 2020-04-25
        • 2022-08-18
        • 1970-01-01
        • 2013-04-12
        相关资源
        最近更新 更多