【问题标题】:How to do an insertion of text before a multi-line regex using sed or awk?如何使用 sed 或 awk 在多行正则表达式之前插入文本?
【发布时间】:2018-07-01 03:01:12
【问题描述】:

给定以下输入(不是字面意思,而是用一些元符号显示):

... any content can be above the match ...
# ... optional comment above the match ...
  # ... optional comment above the match can have spaces before it ...
 "<key>": ... any content can follow ...
... any content can be below the match ...

匹配是^\s*"&lt;key&gt;":,其中&lt;key&gt; 是实际字符串的占位符。请注意,cmets 与 ^\s*#.* 匹配。

我想在匹配的&lt;key&gt; 之前和匹配的&lt;key&gt; 正上方的任何 cmets 之前插入一个文本字符串。可能有可变数量的 cmets,或者根本没有。

我想出了一个使用 sed 的解决方案;但是,它非常难看,因为它使用了tr hack。我希望使用 sed 或 awk 获得更简单的解决方案。

首先,这是一个测试用例:

test.txt:

{
# 1a
# 2a
"key1": true,

# 1b
# 2b
"key2": false,
}

现在我目前的解决方案涉及 sed 并将所有换行符转换为分隔符 ($'\x01'),以便更轻松地进行多行操作。我的示例涉及一个正则表达式,它匹配多个注释行,后跟一个键值对。

# The string to insert before the match
s='# 1x
# 2x
"keyx": null,

'

# Define the key before which to do the insertion:
Key='key2'

# Normalize that string: s -> ns
ns="$(printf '%s' "$s" | tr '\n' $'\x01')"

# Normalize test.txt
tr '\n' $'\x01' < test.txt |
# Perform the multi-line insertion
sed "s/\(^\|\x01\)\(\(\s*#[^\x01]*\x01\)*\)\(\s*\"$Key\":\)/\1$ns\2\4/" |
# Return to standard form with newlines
tr $'\x01' '\n'

使用 test.txt 输入执行上述代码会产生正确且预期的输出:

{
# 1a
# 2a
"key1": true,

# 1x
# 2x
"keyx": null,

# 1b
# 2b
"key2": false,
}

我如何使用 sed 或 awk 改进我在上面所做的工作以生成更易于维护的代码?具体来说:

  • 是否有另一种方法可以使用 sed 而不使用上述 tr hack?
  • 有没有更简单的方法可以使用 awk 做到这一点?

【问题讨论】:

  • @EdMorton 我的假设是这种事情可以在 awk 中更简单地完成。而且,也许有更好的方法来使用 sed 而不依赖于 tr hack。我正在寻找一个更简单的解决方案。我现在的解决方案有效,但其他人很难理解。
  • 嗨 Ed... 我已将 OP 的文本更新为“产生正确且预期的输出”。您对我的方法的确认正是我希望在这里找到更好的方法的原因,以便我可以从现在开始更好地做这样的事情。谢谢。
  • 我的第一句话是:我希望能够在多行正则表达式之前插入一段文本。它可以改进为: ... 在多行正则表达式指示的匹配之前插入一段文本。这是你的意思吗?
  • 我已经尝试按照你上面提到的格式。

标签: bash shell awk sed


【解决方案1】:

在您更新输入可能不包含或包含不同数量的 cmets 之后,这就是编辑(由于编辑它的一些问题,我不得不编辑 v1,所以如果您想要它回来发表评论。 )

sed 真的不做循环或 if/else,只是标签和分支,所以尝试选择一系列行似乎有点复杂。或者至少就我的知识水平而言。

export key='key2'
s='# 1x\n# 2x\n"keyx": null,\n'
key_pattern='[[:space:]]*"'"$key"'":'

sed -n '

/'"$key_pattern"'/ {
  :b; i\
'"$s"'
  p; d  
}

/^[[:space:]]*#/ { 
  h; :a; n; H
  /^[[:space:]]*#/ ba
  /'"$key_pattern"'/ { x; bb; }
  x; p; d; 
}

p
'

这个脚本分为三种类型的模式; key_pattern 匹配但独立的位置(之前没有 cmets):

/'"$key_pattern"'/ {  # here :b creates label b, 
  :b; i\              # and inserts
'"$s"'                # the contents of this line
  p; d                # print then delete from buffer and start next line
}

当一组cmets后面跟着key_pattern时:

/^[[:space:]]*#/ {   # if comment found
  h;                 # copy pattern space into hold space
  :a;                # create label a 
  n; H               # get next line, append to hold space.
  /^[[:space:]]*#/ ba              # if new line is comment, goto `a`
  /'"$key_pattern"'/ { x; bb; }    # else if our pattern retrieve hold 
                                   # and goto `b`
  x; p; d;                         # retrieve hold space, print and delete
}

最后,当该行与其他任何内容都不匹配时:

p;    # print line and start next.

【讨论】:

  • 上述脚本是否处理任意数量的注释行(0 或更多)? cmets 是可选的。
  • @SteveAmerige 我认为应该这样做。
【解决方案2】:

以下代码带有这些假设:

  1. 键和数据之间有空行

  2. 大括号不在别处

    awk '/key2/{$0 = "# 1x\n# 2x\n\"keyx\": null,\n\n"$0}ORS = RT' RS='[{}\n]\n ' 输入文件

这里的主要重点是设置 RS 值,以便分隔每条记录

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2013-12-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-03-02
    • 2018-06-10
    相关资源
    最近更新 更多