【问题标题】:sed or awk script to substitute the structure of a text filesed 或 awk 脚本来替换文本文件的结构
【发布时间】:2015-04-24 15:58:03
【问题描述】:

我想创建一个 sed 或 awk 脚本,在 awk -f script.awk oldfile > newfile 上将给定的文本文件 oldfile 转换为内容

Some Heading
example text

Another Heading
1. example list item, but it
spans over multiple lines
2. list item

进入一个新的文本文件newfile,内容如下:

{Some Heading:} {example text}

{Another Heading:} {
  [item] example list item, but it spans over multiple lines
  [item] list item
}

进一步描述以消除可能的歧义:

  • 脚本应相应地替换每个块(即行,由空行封装)。
  • 在一个文本文件中,可能会出现多个这样的块,但不清楚它们出现的顺序。
  • 脚本应根据标题(即块的第一行)后跟项目列表(以“1.”开头的行表示)有条件地进行替换。
  • 块总是用空行分隔。

如何使用 sed 或 awk 完成此任务? (我使用 zsh 以防万一。)


补充:我刚刚发现我确实需要事先知道块是否是列表:

heading
1. foo
2. bar

到

{list: heading}{
 [item] foo
 [item] bar
}

所以如果是列表,我需要输入“列表:”。这个也可以吗?

【问题讨论】:

    标签: regex shell awk sed


    【解决方案1】:

    使用 awk,您可以执行以下操作:

    awk '/^$/ { print block (list ? "\n}" : "}"); block = ""; next } block == "" { block = "{" $0 ":} {"; list = 0; next } /^[0-9]+\. / { list = 1; sub(/^[0-9]+\. /, ""); block = block "\n  [item] " $0; next } { block = block (list ? " " : "") $0 } END { print block (list ? "\n}" : "}") }' filename
    

    代码在哪里:

    #!/usr/bin/awk -f
    
    /^$/ {                               # empty line: print converted block
      print block (list ? "\n}" : "}")   # Whether there's a newline before the
      block = ""                         # closing } depends on whether this is
      next                               # a list. Reset block buffer.
    }
    block == "" {                        # in the first line of a block:
      block = "{" $0 ":} {"              # format header
      list = 0                           # reset list flag
      next
    }
    /^[0-9]+\. / {                       # if a data line opens a list
      list = 1                           # set list flag
      sub(/^[0-9]+\. /, "")              # remove number
      block = block "\n  [item] " $0     # format line
      next
    }
    {                                    # if it doesn't, just append it. Space
      block = block (list ? " " : "") $0 # inside a list to not fuse words.
    }
    END {                                # and at the very end, print the last
      print block (list ? "\n}" : "}")   # block
    }
    

    sed 也可以,但更难阅读:

    #!/bin/sed -nf
    
    /^$/ {                       # empty line: print converted block
      x                          # fetch it from the hold buffer
      s/$/}/                     # append closing }
      /\n  \[item\]/ s/}$/\n}/   # in a list, put in a newline before it
      p                          # print
      d                          # and we're done here. Hold buffer is now empty.
    }
    x                            # otherwise: inspect the hold buffer
    // {                         # if it is empty (reusing last regex)
      x                          # get back the pattern space
      s/.*/{&:}{/                # Format header
      h                          # hold it.
      d                          # we're done here.
    }
    x                            # otherwise, get back the pattern space
    /^[0-9]\+\. / {              # if the line opens a list
      s///                       # remove the number (reusing regex)
      s/.*/  [item] &/           # format the line
      H                          # append it to the hold buffer.
      ${                         # if it is the last line
        s/.*/}/                  # append a closing bracket
        H                        # to the hold buffer
        x                        # swap it with the hold buffer
        p                        # and print that.
      }
      d                          # we're done.
    }
                                 # otherwise (not opening a list item)
    H                            # append line to the hold buffer
    x                            # fetch back the hold buffer to work on it
    
    /\n  \[item\]/ {             # if we're in a list
      s/\(.*\)\n/\1 /            # replace the last newline (that we just put there)
                                 # with a space
      ${
        s/$/\n}/                 # if this is the last line, append \n}
        p                        # and print
      }
      x                          # put the half-assembled block in the hold buffer
      d                          # and we're done
    }
    s/\(.*\)\n/\1/               # otherwise (not in a list): just remove the newline
    ${
      s/$/}/                     # if this is the last line, append closing bracket
      p                          # print
    }
    x                            # put half-assembled block in the hold buffer.
    

    【讨论】:

    • 喜欢轻描淡写的...rather more difficult to read :-)。
    【解决方案2】:

    sed 是面向行的,因此最适合在单行上进行简单替换。

    只需在段落模式下使用 awk (RS=""),因此每个空行分隔文本块都被视为一条记录,并将每个段落中的每一行视为记录的一个字段 (FS="\n"):

    $ cat tst.awk
    BEGIN { RS=""; ORS="\n\n"; FS="\n" }
    {
        printf "{" (/\n[0-9]+\./ ? "list: %s" : "%s:") "} {", $1
        inList = 0
        for (i=2; i<=NF; i++) {
            if ( sub(/^[0-9]+\./,"  [item]",$i) ) {
                printf "\n"
                inList = 1
            }
            else if (inList) {
                printf " "
            }
            printf "%s", $i
        }
        print (inList ? "\n" : "") "}"
    }
    $
    $ awk -f tst.awk file
    {Some Heading:} {example text}
    
    {list: Another Heading} {
      [item] example list item, but it spans over multiple lines
      [item] list item
    }
    

    【讨论】:

    • 我更新了我的答案。请注意,在执行标题行的 printf 时,只需进行最细微的调整即可检查记录中是否存在换行符和数字。我想知道在 sed 脚本中需要什么...... :-)。关键是 - 不要将 sed 用于任何涉及处理多行的事情,在 1970 年代中期,当 awk 被发明时,所有 seds 神秘的语言结构都变得过时了,如今人们只是为了挑战而这样做,比如解决一个复杂的填字游戏。
    【解决方案3】:

    另一个 awk 版本(类似于 Eds)

    BEGIN{RS="";FS="\n"}
    {
        {printf "%s", "{"(/\n[0-9]+\./?"Line: ":"")$1":} {"
        for(i=2;i<=NF;i++)
        printf "%s",sub(/^[0-9]+\./,"  [item]",$i)&&++x?"\n"$i:$i
        print x?"\n}":"}""\n"
        x=0
    }
    

    输出

    $awk -f test.awk file
    
    {Some Heading:} {example text}
    
    {Another Heading:} {
      [item] example list item, but itspans over multiple lines
      [item] list item
    }
    

    工作原理

    BEGIN{RS="";FS="\n"}
    

    以空行分隔的块形式读取记录。
    将字段读取为行。

    {printf "%s", "{"(/\n[0-9]+\./?"List: ":"")$1":} {"
    

    以指定格式打印第一个字段(行),注意 printf 用于省略换行符。 检查记录的任何部分是否包含换行符,然后是数字和句点,如果包含则添加列表。

    for(i=2;i<=NF;i++)
    

    从第二个字段循环到最后一个字段。 NF 是字段数。

    我会分开下一点。

    printf "%s"
    

    打印一个字符串,printf再次用来控制换行

    sub(/^[0-9]+\./,"  [item]",$i)&&++x?"\n"$i:$i
    

    这实际上是一个使用三元运算符a?b:c 的 if else 语句。 如果无法完成 sub 将返回 0 并且 x 不会递增,因此该行将按原样打印。
    如果 sub 成功,它将用[item] 替换该行开头的数字,增加 x 并在新行之前打印新行。

    print x?"\n}":"}""\n"
    

    再次使用三元运算符来检查 x 是否增加。如果它在 } 之前打印一个换行符,否则只是 rpints }。为记录之间的双换行符打印一个换行符。

    【讨论】:

    • 记得在记录之间重置x,你需要预先添加一个测试来打印list:。
    • @EdMorton Yep 错过了 x 因为它只有两行输入谢谢!也没有看到列表
    • 他们问题的第二部分:Addition: I just found out....
    • 感谢您的解释。它确实为其他答案增添了一些东西。
    猜你喜欢
    • 2018-07-02
    • 1970-01-01
    • 1970-01-01
    • 2018-08-20
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多