【问题标题】:print columns based on column name根据列名打印列
【发布时间】:2021-04-21 00:54:28
【问题描述】:

假设我有一个文件test.txt,其中包含

a,b,c,d,e
1,2,3,4,5
6,7,8,9,10

我想根据匹配的列名打印出来自另一个文本文件或数组的列。因此,例如,如果给我

arr=(a b c)

我希望我的输出是

a,b,c
1,2,3
6,7,8

如何使用 bash 实用程序/awk/sed 执行此操作?我的实际文本文件是 3GB(我想要匹配列值的行实际上是第 3 行),因此非常感谢有效的解决方案。这是我目前所拥有的:

for j in "${arr[@]}"; do awk -F ',' -v a=$j '{ for(i=1;i<=NF;i++) {if($i==a) {print $i}}}' test.txt; done

但我得到的输出是

a
b
c

这不仅缺少其他行,而且每个列名都打印在一行上。

【问题讨论】:

    标签: bash awk sed


    【解决方案1】:

    使用您展示的示例,请尝试以下操作。代码正在读取 2 个文件 another_file.txt(根据样本有 a b c)和名为 test.txt 的实际 Input_file(其中包含所有值)。

    awk '
    FNR==NR{
      for(i=1;i<=NF;i++){
        arr[$i]
      }
      next
    }
    FNR==1{
      for(i=1;i<=NF;i++){
        if($i in arr){
          valArr[i]
          header=(header?header OFS:"")$i
        }
      }
      print header
      next
    }
    {
      val=""
      for(i=1;i<=NF;i++){
        if(i in valArr){
           val=(val?val OFS:"")$i
        }
      }
      print val
    }
    ' another_file.txt FS="," OFS="," test.txt
    

    输出如下:

    a,b,c
    1,2,3
    6,7,8
    

    说明:为上述解决方案添加详细说明。

    awk '                                         ##Starting awk program from here.
    FNR==NR{                                      ##Checking condition which will be TRUE while reading another_text file here.
      for(i=1;i<=NF;i++){                         ##Traversing through all fields of current line.
        arr[$i]                                   ##Creating arr with index of current field value.
      }
      next                                        ##next will skip all statements from here.
    }
    FNR==1{                                       ##Checking if this is 1st line for test.txt file.
      for(i=1;i<=NF;i++){                         ##Traversing through all fields of current line.
        if($i in arr){                            ##If current field values comes in arr then do following.
          valArr[i]                               ##Creating valArr which has index of current field number.
          header=(header?header OFS:"")$i         ##Creating header which has each field value in it.
        }
      }
      print header                                ##Printing header here.
      next                                        ##next will skip all statements from here.
    }
    {
      val=""                                      ##Nullifying val here.
      for(i=1;i<=NF;i++){                         ##Traversing through all fields of current line.
        if(i in valArr){                          ##Checking if i is present in valArr then do following.
           val=(val?val OFS:"")$i                 ##Creating val which has current field value.
        }
      }
      print val                                   ##printing val here.
    }
    ' another_file.txt FS="," OFS="," test.txt    ##Mentioning Input_file names here.
    

    【讨论】:

    • 感谢您的解释和解释,解决方案完美运行。
    【解决方案2】:

    您可以通过以下方式在单遍 awk 命令中执行此操作:

    arr=(a c e)
    awk -v cols="${arr[*]}" 'BEGIN {FS=OFS=","; n=split(cols, tmp, / /); for (i=1; i<=n; ++i) hdr[tmp[i]]} NR==1 {for (i=1; i<=NF; ++i) if ($i in hdr) hnum[i]} {for (i=1; i<=NF; ++i) if (i in hnum) {printf "%s%s", (f ? OFS : ""), $i; f=1} f=0; print ""}' file
    
    a,c,e
    1,3,5
    6,8,10
    

    更易读的形式:

    awk -v cols="${arr[*]}" '
    BEGIN {
       FS = OFS = ","
       n = split(cols, tmp, / /)
       for (i=1; i<=n; ++i)
          hdr[tmp[i]]
    }
    NR == 1 {
       for (i=1; i<=NF; ++i)
          if ($i in hdr)
             hnum[i]
    }
    {
       for (i=1; i<=NF; ++i)
          if (i in hnum) {
             printf "%s%s", (f ? OFS : ""), $i
             f = 1
          }
          f = 0
          print ""
    }' file
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-12-06
      • 1970-01-01
      • 2013-10-22
      • 2018-01-07
      • 1970-01-01
      • 2015-06-15
      相关资源
      最近更新 更多