【问题标题】:divide each column by max value/last value将每一列除以最大值/最后一个值
【发布时间】:2020-11-30 05:24:10
【问题描述】:

我有一个这样的矩阵:

A   25  27  50

B   35  37  475

C   75  78  80

D   99  88  76

0   234 230 681

最后一行是列中所有元素的总和——也是最大值。

我想得到一个矩阵,其中每个值除以列中的最后一个值(例如,对于第 2 列中的第一个数字,我想要“25/234 =”):

A   0.106837606837607   0.117391304347826   0.073421439060206

B   0.14957264957265    0.160869565217391   0.697503671071953

C   0.320512820512821   0.339130434782609   0.117474302496329

D   0.423076923076923   0.382608695652174   0.11160058737151

另一个线程中的答案对一列给出了可接受的结果,但我无法在所有列上循环它。

$ awk 'FNR==NR{max=($2+0>max)?$2:max;next} {print $1,$2/max}' file file

(这里提供了这个答案:normalize column data with maximum value of that column)

如果有任何帮助,我将不胜感激!

【问题讨论】:

    标签: matrix awk


    【解决方案1】:

    除了@RavinderSingh13 的出色方法之外,您还可以使用例如隔离输入文件中的最后一行tail -n1 Input_file 然后在BEGIN 规则中使用split() 命令来分隔值。然后,您可以使用awk 对文件进行单次传递,以按照您的指示更新值。最后,您可以将输出通过管道传输到head -n-1 以删除不需要的最后一行,例如

    awk -v lline="$(tail -n1 Input_file)" '
        BEGIN { split(lline,a," ") }
        {
            printf "%s", $1
            for(i=2; i<=NF; i++)
                printf "  %.15lf", $i/a[i]
            print ""
        }
    ' Input_file | head -n-1
    

    使用/输出示例

    $ awk -v lline="$(tail -n1 Input_file)" '
    >     BEGIN { split(lline,a," ") }
    >     {
    >         printf "%s", $1
    >         for(i=2; i<=NF; i++)
    >             printf "  %.15lf", $i/a[i]
    >         print ""
    >     }
    > ' Input_file | head -n-1
    A  0.106837606837607  0.117391304347826  0.073421439060206
    B  0.149572649572650  0.160869565217391  0.697503671071953
    C  0.320512820512821  0.339130434782609  0.117474302496329
    D  0.423076923076923  0.382608695652174  0.111600587371512
    

    (注意:这假定您的文件中没有尾随空行,并且每行之间确实没有空行。如果有,请告诉我)

    这些方法之间的差异在很大程度上可以忽略不计。在每种情况下,您总共要通过文件 3 次。这里是tail、awk,然后是head。在另一种情况下使用wc,然后使用awk 两次通过。

    如果您有任何问题,请告诉我们。

    【讨论】:

      【解决方案2】:

      第一个解决方案:您能否尝试在 GNU awk 中使用所示示例进行跟踪、编写和测试。根据 OP 显示的样本,精确的 15 个浮点数:

      awk -v lines=$(wc -l < Input_file) '
      FNR==NR{
        if(FNR==lines){
          for(i=2;i<=NF;i++){ arr[i]=$i }
        }
        next
      }
      FNR<lines{
        for(i=2;i<=NF;i++){ $i=sprintf("%0.15f",(arr[i]?$i/arr[i]:"NaN")) }
        print
      }
      ' Input_file  Input_file
      

      第二个解决方案:如果您不关心浮点是特定点,请尝试以下操作。

      awk -v lines=$(wc -l < Input_file) '
      FNR==NR && FNR==lines{
        for(i=2;i<=NF;i++){ arr[i]=$i }
        next
      }
      FNR<lines && FNR!=NR{
        for(i=2;i<=NF;i++){ $i=(arr[i]?$i/arr[i]:"NaN") }
        print
      }
      ' Input_file Input_file
      

      OR(将FNR==lines 放置在FNR==NR 条件中):

      awk -v lines=$(wc -l < Input_file) '
      FNR==NR{
        if(FNR==lines){
          for(i=2;i<=NF;i++){ arr[i]=$i }
        }
        next
      }
      FNR<lines{
        for(i=2;i<=NF;i++){ $i=(arr[i]?$i/arr[i]:"NaN") }
        print
      }
      ' Input_file  Input_file
      

      说明:为上述添加详细说明。

      awk -v lines=$(wc -l < Input_file) '         ##Starting awk program from here, creating lines which variable which has total number of lines in Input_file here.
      FNR==NR{                                     ##Checking condition FNR==NR which will be TRUE when first time Input_file is being read.
        if(FNR==lines){                            ##Checking if FNR is equal to lines then do following.
          for(i=2;i<=NF;i++){ arr[i]=$i }          ##Traversing through all fields here of current line and creating an array arr with index of i and value of current field value.
        }
        next                                       ##next will skip all further statements from here.
      }
      FNR<lines{                                   ##Checking condition if current line number is lesser than lines, this will execute when 2nd time Input_file is being read.
        for(i=2;i<=NF;i++){ $i=sprintf("%0.15f",(arr[i]?$i/arr[i]:"NaN")) } ##Traversing through all fields here and saving value of divide of current field with arr current field value with 15 floating points into current field.
        print                                      ##Printing current line here.
      }
      ' Input_file  Input_file                     ##Mentioning Input_file names here.
      

      【讨论】:

      • 将在几分钟左右添加详细说明。
      • 所有三种解决方案似乎都运行良好 - 谢谢!明天我会做一些详细的测试。我期待着从你的解释中学习。
      • @steff2j,当然(添加解释在一分钟左右),你可以看到这个链接What one could do when someone gets helpful answer on SO欢呼,快乐学习。
      • lines=$(tail -n1 Input_file) 然后split (line, ...) in begin :) 怎么样
      猜你喜欢
      • 1970-01-01
      • 2015-01-09
      • 1970-01-01
      • 2015-05-30
      • 1970-01-01
      • 1970-01-01
      • 2014-04-06
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多