【问题标题】:Save modifications in place with NON GNU awk使用非 GNU awk 保存修改
【发布时间】:2020-04-02 05:10:38
【问题描述】:

我遇到了一个问题(关于 SO 本身),OP 必须对 Input_file(s) 本身进行编辑和保存操作。

我知道对于单个 Input_file 我们可以执行以下操作:

awk '{print "test here..new line for saving.."}' Input_file > temp && mv temp Input_file

现在假设我们需要更改相同格式的文件(假设为 .txt)。

我对这个问题的尝试/想法:它的方法是通过 .txt 文件的 for 循环并调用单个 awk 是痛苦的,而不是推荐的过程,因为它会浪费不必要的 cpu 周期,并且对于更多数量的文件会更慢。

那么在这里可以做些什么来使用不支持就地选项的非 GNU awk 对多个文件执行就地编辑。我也经历了这个线程Save modifications in place with awk,但是对于非 GNU awk 的副手和在 awk 本身内就地更改多个文件没有什么意义,因为非 GNU awk 不会有 inplace 选项。

注意: 为什么我要添加 bash 标记,因为在我的回答部分中,我使用 bash 命令将临时文件重命名为它们的实际 Input_file 名称,所以添加它.



编辑:根据 Ed sir 的评论,在此处添加示例示例,尽管此线程代码的目的也可以用于通用目的的就地编辑。

输入文件示例:

cat test1.txt
onetwo three
tets testtest

cat test2.txt
onetwo three
tets testtest

cat test3.txt
onetwo three
tets testtest

预期输出示例:

cat test1.txt
1
2

cat test2.txt
1
2

cat test3.txt
1
2

【问题讨论】:

  • 有趣且中肯的awk问题++

标签: linux bash shell awk inplace-editing


【解决方案1】:

由于这个线程的主要目的是如何在非 GNU awk 中进行就地保存,所以我首先发布它的模板,这将帮助任何有任何要求的人,他们需要添加/附加 BEGIN 和 @987654323 @ 部分在他们的代码中保持他们的主要块按照他们的要求,然后它应该进行就地编辑:

注意: 以下会将其所有输出写入 output_file,因此如果您想将任何内容打印到标准输出,请仅添加 print... 语句而不添加 > (out)在下面。

通用模板:

awk -v out_file="out" '
FNR==1{
close(out)
out=out_file count++
rename=(rename?rename ORS:"") "mv \047" out "\047 \047" FILENAME "\047"
}
{
    .....your main block code.....
}
END{
 if(rename){
   system(rename)
 }
}
' *.txt


提供的具体示例解决方案:

我在awk 本身中提出了以下方法(对于添加的示例,以下是我解决此问题并将输出保存到 Input_file 本身的方法)

awk -v out_file="out" '
FNR==1{
  close(out)
  out=out_file count++
  rename=(rename?rename ORS:"") "mv \047" out "\047 \047" FILENAME "\047"
}
{
  print FNR > (out)
}
END{
  if(rename){
    system(rename)
  }
}
' *.txt

注意:这只是将编辑后的输出保存到 Input_file(s) 本身的测试,可以在他们的程序中使用它的 BEGIN 部​​分以及它的 END 部分,主要部分应该按照具体问题本身的要求。

公平警告:此外,由于这种方法会在路径中生成一个新的临时输出文件,因此更好地确保我们在系统上有足够的空间,尽管在最终结果中这将保持只有主要的 Input_file(s),但在操作期间它需要系统/目录上的空间



以下是对上述代码的测试。

用一个例子执行程序:让我们假设以下是.txt Input_file(s):

cat << EOF > test1.txt
onetwo three
tets testtest
EOF

cat << EOF > test2.txt
onetwo three
tets testtest
EOF

cat << EOF > test3.txt
onetwo three
tets testtest
EOF

现在当我们运行以下代码时:

awk -v out_file="out" '
FNR==1{
  close(out)
  out=out_file count++
  rename=(rename?rename ORS:"") "mv \047" out "\047 \047" FILENAME "\047"
}
{
  print "new_lines_here...." > (out)
}
END{
  if(rename){
    system("ls -lhtr;" rename)
  }
}
' *.txt

注意:我有意将ls -lhtr 放在system 部分中,以查看它正在创建哪些输出文件(临时),因为稍后它会将它们重命名为他们的真实姓名。

-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test2.txt
-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test1.txt
-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test3.txt
-rw-r--r-- 1 runner runner  38 Dec  9 05:33 out2
-rw-r--r-- 1 runner runner  38 Dec  9 05:33 out1
-rw-r--r-- 1 runner runner  38 Dec  9 05:33 out0

当我们在awk 脚本运行完成后执行ls -lhtr 时,我们只能看到.txt 文件。

-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test2.txt
-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test1.txt
-rw-r--r-- 1 runner runner  27 Dec  9 05:33 test3.txt


说明:在此处添加上述命令的详细说明:

awk -v out_file="out" '                                    ##Starting awk program from here, creating a variable named out_file whose value SHOULD BE a name of files which are NOT present in our current directory. Basically by this name temporary files will be created which will be later renamed to actual files.
FNR==1{                                                    ##Checking condition if this is very first line of current Input_file then do following.
  close(out)                                               ##Using close function of awk here, because we are putting output to temp files and then renaming them so making sure that we shouldn't get too many files opened error by CLOSING it.
  out=out_file count++                                     ##Creating out variable here, whose value is value of variable out_file(defined in awk -v section) then variable count whose value will be keep increment with 1 whenever cursor comes here.
  rename=(rename?rename ORS:"") "mv \047" out "\047 \047" FILENAME "\047"     ##Creating a variable named rename, whose work is to execute commands(rename ones) once we are done with processing all the Input_file(s), this will be executed in END section.
}                                                          ##Closing BLOCK for FNR==1  condition here.
{                                                          ##Starting main BLOCK from here.
  print "new_lines_here...." > (out)                       ##Doing printing in this example to out file.
}                                                          ##Closing main BLOCK here.
END{                                                       ##Starting END block for this specific program here.
  if(rename){                                              ##Checking condition if rename variable is NOT NULL then do following.
    system(rename)                                         ##Using system command and placing renme variable inside which will actually execute mv commands to rename files from out01 etc to Input_file etc.
  }
}                                                          ##Closing END block of this program here.
' *.txt                                                    ##Mentioning Input_file(s) with their extensions here.

【讨论】:

    【解决方案2】:

    shell 解决方案很简单,而且可能足够快:

    for f in *.txt
    do  awk '...' "$f" > "$f.tmp"
        mv "$f.tmp" "$f"
    done
    

    仅当您最终证明这太慢时才搜索不同的解决方案。请记住:过早优化是万恶之源。

    【讨论】:

      【解决方案3】:

      如果我要尝试这样做,我可能会选择这样的东西:

      $ cat ../tst.awk
      FNR==1 { saveChanges() }
      { print FNR > new }
      END { saveChanges() }
      
      function saveChanges(   bak, result, mkBackup, overwriteOrig, rmBackup) {
          if ( new != "" ) {
              bak = old ".bak"
              mkBackup = "cp \047" old "\047 \047" bak "\047; echo \"$?\""
              if ( (mkBackup | getline result) > 0 ) {
                  if (result == 0) {
                      overwriteOrig = "mv \047" new "\047 \047" old "\047; echo \"$?\""
                      if ( (overwriteOrig | getline result) > 0 ) {
                          if (result == 0) {
                              rmBackup = "rm -f \047" bak "\047"
                              system(rmBackup)
                          }
                      }
                  }
              }
              close(rmBackup)
              close(overwriteOrig)
              close(mkBackup)
          }
          old = FILENAME
          new = FILENAME ".new"
      }
      
      $ awk -f ../tst.awk test1.txt test2.txt test3.txt
      

      我宁愿先将原始文件复制到备份中,然后对原始文件进行保存更改,但这样做会更改每个输入文件的 FILENAME 变量的值,这是不可取的。

      请注意,如果您的目录中有一个名为 whatever.bak 或 whatever.new 的原始文件,那么您将用临时文件覆盖它们,因此您也需要为此添加一个测试。调用mktemp 来获取临时文件名会更可靠。

      在这种情况下更有用的东西是执行任何其他命令并执行“就地”编辑部分的工具,因为它可用于为 POSIX sed、awk、grep、tr 提供“就地”编辑,无论何时都不需要您将脚本的语法更改为print &gt; out 等,每次您想打印一个值时。一个简单易碎的例子:

      $ cat inedit
      #!/bin/env bash
      
      for (( pos=$#; pos>1; pos-- )); do
          if [[ -f "${!pos}" ]]; then
              filesStartPos="$pos"
          else
              break
          fi
      done
      
      files=()
      cmd=()
      for (( pos=1; pos<=$#; pos++)); do
          arg="${!pos}"
          if (( pos < filesStartPos )); then
              cmd+=( "$arg" )
          else
              files+=( "$arg" )
          fi
      done
      
      tmp=$(mktemp)
      trap 'rm -f "$tmp"; exit' 0
      
      for file in "${files[@]}"; do
          "${cmd[@]}" "$file" > "$tmp" && mv -- "$tmp" "$file"
      done
      

      你会使用如下:

      $ awk '{print FNR}' test1.txt test2.txt test3.txt
      1
      2
      1
      2
      1
      2
      
      $ ./inedit awk '{print FNR}' test1.txt test2.txt test3.txt
      
      $ tail test1.txt test2.txt test3.txt
      ==> test1.txt <==
      1
      2
      
      ==> test2.txt <==
      1
      2
      
      ==> test3.txt <==
      1
      2
      

      inedit 脚本的一个明显问题是当您有多个输入文件时,难以将输入/输出文件与命令分开识别。上面的脚本假定所有输入文件在命令末尾显示为一个列表,并且命令一次对它们运行一个,但当然这意味着您不能将它用于需要 2 个或更多文件的脚本一次,例如:

      awk 'NR==FNR{a[$1];next} $1 in a' file1 file2
      

      或在 arg 列表中的文件之间设置变量的脚本,例如:

      awk '{print $7}' FS=',' file1 FS=':' file2
      

      让它更健壮留给读者作为练习,但请以xargs 概要作为起点,了解健壮的inedit 需要如何工作:-)。

      【讨论】:

        猜你喜欢
        • 2021-04-23
        • 1970-01-01
        • 2012-11-23
        • 2018-08-25
        • 1970-01-01
        • 1970-01-01
        • 2018-12-30
        相关资源
        最近更新 更多