【问题标题】:AWK: operations on multiple-columns data fillesAWK:对多列数据填充的操作
【发布时间】:2021-04-27 08:07:06
【问题描述】:

我正在通过集成到 bash 脚本中的以下 AWK 代码(进行所有统计计算)(与数据文件一起操作)来处理对多列格式的大量数据填充的分析:

#!/bin/bash
home="$PWD"
# folder with the outputs
rescore="${home}"/rescore 
# folder with the folders to analyse
storage="${home}"/results
#cd "${home}"/results
cd ${storage}
csv_pattern='*_filt.csv'


while read -r d; do
awk '
FNR==1 {
   if (n) {                     # calculate the results of previous file
      m = s / n                 # mean
      mean[suffix] = m          # store the mean in an array
      lowest[suffix] = min      # lowest value of dG - correspond to the upper number in the original CSV
   }
   prefix=suffix=FILENAME
   sub(/_.*/, "", prefix)
   sub(/\/[^\/]+$/, "", suffix)
   sub(/^.*_/, "", suffix)
   s = 0                        # sum of $3
   s2 = 0                       # sum of $3 ** 2
   n = 0                        # count of samples
   min = 0                      # highest value of $3
}
FNR > 1 {
   s += $3
   s2 += $3 * $3
   ++n
   if ($3 < min) min = $3       # update the lowest value
}
END {
  if (n) {                     # just to avoid division by zero
   m = s / n
   lowest[suffix] = min
  }
   print "Lig(CNE)", "dG(mean)", "dG(min)"
   for (i in mean)
      printf "%s %.2f %.2f %.2f\n", i, mean[i], lowest[i]
}'  "${d}_"*/${str} > "${rescore}/${str_name}/"${d%%_*}".csv"
done < <(find . -maxdepth 1 -type d -name '*_*_*' | awk -F '[_/]' '!seen[$2]++ {print $2}')

基本上,在循环运行时,脚本会为每个 CSV 文件计算第三列 (dG) 中数字的平均值以及检测其最小值(始终对应于 ID=1 的行) :

# input *_filt.csv located in the folder 10V1_cne_lig1001
ID, POP, dG
1, 142, -5.6500 # this is dG min
2, 10, -5.5000
3, 2, -4.9500
4, 150, -4.1200 # this is pop(MAX)

并将结果保存在另一个多列输出文件中(对于 10 个已处理的 CSV,它计为 10 行),包含每个已处理的 CSV 名称的一部分(对应的前缀用作行的 ID),其 dG (平均值)和 dG(最小值):

# output.csv
Lig(CNE) dG(mean) dG(min)
lig1 -6.78 -7.23
lig2 -5.56 -5.76
lig3 -7.30 -8.69
lig4 -7.98 -8.60
lig5 -6.78 -7.16
lig6 -6.24 -6.50
lig7 -7.44 -8.01
lig8 -4.62 -5.60
lig9 -7.26 -7.48
lig10 -5.9 -6.03

我需要在代码的 AWK 部分中添加一种可能性,以检测并在单个列中打印 $3 (dG) 中的值,这些值将在初始 csv 的 $2(列 pop )中具有最大值。在上面的例子中,这个 dG 的值是 -4.1200,它对应于 CSV 的第 4 行,基于在第二列中检测到的最高数字 (150)。因此,目标是将第四列打印到 output.csv,其中将包含与 $2 (pop) 中的最大值相对应的 $3 (dG) 值。

【问题讨论】:

    标签: bash awk multiple-columns


    【解决方案1】:

    至于awk部分,请您试试:

    awk -F ", *" '                  # set field separator to comma, followed by 0 or more whitespaces
    FNR==1 {
       if (n) {                     # calculate the results of previous file
          m = s / n                 # mean
          mean[suffix] = m          # store the mean in an array
          lowest[suffix] = min      # lowest value of dG - correspond to the upper number in the original CSV
          highest[suffix] = fourth  # dG of highest pop
       }
       prefix=suffix=FILENAME
       sub(/_.*/, "", prefix)
       sub(/\/[^\/]+$/, "", suffix)
       sub(/^.*_/, "", suffix)
       s = 0                        # sum of $3
       s2 = 0                       # sum of $3 ** 2
       n = 0                        # count of samples
       min = 0                      # lowest value of $3 (assuming all $3 < 0)
       max = 0                      # highest value of $2 (assuming all $2 > 0)
    }
    FNR > 1 {
       s += $3
       s2 += $3 * $3
       ++n
       if ($3 < min) min = $3       # update the lowest value
       if ($2 > max) {
          max = $2                  # update the highest value
          fourth = $3               # to be printed in the fourth column
       }
    }
    END {
       if (n) {                     # just to avoid division by zero
          m = s / n
          mean[suffix] = m          # store the mean in an array
          lowest[suffix] = min
          highest[suffix] = fourth  # dG of highest pop
       }
       print "Lig(CNE)", "dG(mean)", "dG(min)", "dG(highest pop)"
       for (i in mean)
          printf "%s %.2f %.2f %.2f\n", i, mean[i], lowest[i], highest[i]
    }' *_filt.csv
    
    • 将字段分隔符设置为, 很重要。否则数值比较可能会中断。
    • 我保留了一些未使用的变量(例如 s2),它们可能会用于您未来的更新计划。

    【讨论】:

    • 非常感谢!只有一个问题会“将字段分隔符设置为 ”,会影响以 3 美元进行的其他一些计算(例如平均值、rmsd 等的计算)?
    • 简短的回答是no。由于今天正数之间的比较,可能的问题已经变得明显。如果不将字段分隔符分配给“,”,则字段将按默认字段分隔符、空白字符进行拆分,然后每个字段可能包含尾随逗号,例如10, 和2,。大多数算术运算(包括负数之间的比较)都会自动去掉逗号。
    • 10,和2,等正数比较时出现问题。由于它们看起来像字符串而不是数字,awk 尝试在它们之间执行字符串比较,从而产生2, is greater than 10,。我在之前的回答中错过了潜在的问题。顺便说一句,我已将选项 -F, 的另一个修改添加到 -F ", *" 中,这将删除字段的前导空格,以防万一。
    • 好的,非常感谢!我将运行此更新的简短基准
    • 刚刚完成了 > 2000 csv 填充的测试,因此使用 awk -F ", *" ' 或 awk -F, ' 基本上没有区别,因此表明两个版本都可以完美运行!干杯
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-10-12
    • 1970-01-01
    • 1970-01-01
    • 2020-04-27
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多