【问题标题】:Adding the portion of a line that comes between the first bracket and the first question mark to the end of the line将第一个括号和第一个问号之间的行的部分添加到行尾
【发布时间】:2020-12-05 00:31:57
【问题描述】:

我正在尝试在文件中添加从每行开头到行尾的部分。目前文件格式如下:

1.1) This is a sample question? Yes it is a sample question
1.2) Are you quite sure it is a sample question? I am quite sure
...

我想要做的是将每行开头的问题(而不是数字)添加到行尾,本质上是制作一个格式如下的文件:

1.1) This is a sample question? Yes it is a sample question This is a sample question
1.2) Are you quite sure it is a sample question? I am quite sure Are you quite sure it is a sample question
...

我已经对原始文本文件进行了相当多的重组,包括删除除相关问题末尾的问号之外的所有问号以及除末尾的那些之外的所有右括号每行的编号。

我在这里的理由是使用右括号作为标记,指示要重复的部分从哪里开始,而问号作为标记,表示要重复的部分在哪里结束。但是,在实际尝试实现这一点时,我已经干了。

我假设我需要使用遍历每一行的for 循环,当它看到) 时激活,然后将每个空格分隔的字符添加到行尾,直到它看到?,在这一点上它停止并移动到下一行;但是,我在 bash 中努力实现这一点。

【问题讨论】:

  • 提示:正则表达式

标签: bash for-loop awk sed


【解决方案1】:

假设:

  • 所有感兴趣的行的格式为:<text> + ) + <text2copy> + ? + <more_text
  • 对于我们想要附加到行尾的感兴趣的行:<space> + <text2copy>
  • 所有其他行都不要管

样本数据:

$ cat questions.dat
1.1) This is a sample question? Yes it is a sample question
ignore this line and do nothing to it
1.2) Are you quite sure it is a sample question? I am quite sure

一个想法使用awk:

awk '
/).*\?/ { split($0,arr,"[)?]")   # if line contains ")" + <text> + "?" then split
                                 # the line using ")" and "?" as delimiters, placing 
                                 # results into array "arr[]"
          $0 = $0 arr[2]         # append 2nd element of array to end of line
        }
1                                # print current line
' questions.dat

以上生成:

1.1) This is a sample question? Yes it is a sample question This is a sample question
ignore this line and do nothing to it
1.2) Are you quite sure it is a sample question? I am quite sure Are you quite sure it is a sample question

另一个使用sed 和捕获组的想法:

$ sed -E 's/^[^)]*[)] ([^?]*)[?].*/& \1/' questions.dat

地点:

  • -E - 启用扩展的正则表达式支持
  • ^[^)]*[)] - 匹配行首 (^) + 一些不包括 ) + ) + &lt;space&gt; 的字符
  • ([^?*) - [1st capture group] 匹配所有内容,但不包括 ?
  • [?].* - 匹配从 ? 到行尾
  • &amp; \1 - 打印我们的正则表达式匹配(在这种情况下是整行)+ &lt;space&gt; + 1st capture group

以上生成:

1.1) This is a sample question? Yes it is a sample question This is a sample question
ignore this line and do nothing to it
1.2) Are you quite sure it is a sample question? I am quite sure Are you quite sure it is a sample question

【讨论】:

  • 在诸如 awk 之类的 ERE 中,使用 /).*?/ 是未定义的行为,因为 ? 是一个正则表达式重复元字符,并且根据 POSIX,在重复元字符之后它是未定义的重复元字符(* 在这种情况下) 方法。 ITYM/).*\?/。你不需要在 sed 中转义 ) - 如果前面有 (,它只在一个正则表达式 metachar 上。
  • 按照您的建议更新了awk;至于sed 评论...我要么必须a) 转义) 或b) 将它括起来([)]) 否则sed 会出现sed: -e expression #1, char 28: Unmatched ) or \) - GNU sed 4.4 错误;继续并用括号版本替换了逃生版本
  • 嗯,猜猜这是那个版本的 sed 中的一个错误。每 the POSIX spec for EREs: ) The &lt;right-parenthesis&gt; shall be special when matched with a preceding &lt;left-parenthesis&gt;, both outside a bracket expression. 否则它只是一个文字字符。
  • 我刚刚证实了这一点。 echo 'foo)bar' | awk '{sub(/o)/,"x")}1' 和 sed -E 's/o)/x/' 在 MacOS 上使用 BSD sed 都输出 foxbar,但 GNU sed 4.4 和 4.8 都失败了 sed: -e expression #1, char 7: Unmatched ) or \),正如你所描述的。
【解决方案2】:

我将利用 GNU AWK 接受正则表达式作为字段分隔符的能力。设quest.txt内容为:

1.1) This is a sample question? Yes it is a sample question
1.2) Are you quite sure it is a sample question? I am quite sure

然后

awk 'BEGIN{FS="[)?]"}{print $0$2}' quest.txt

输出:

1.1) This is a sample question? Yes it is a sample question This is a sample question
1.2) Are you quite sure it is a sample question? I am quite sure Are you quite sure it is a sample question

解释:我告诉AWK 将以下任何字符:)? 视为字段分隔符,因此每行分为三个字段。然后我得到整行($0),我将)和?($2)之间的部分连接起来。请注意,) 后面有空格,因此我们不需要包含另一个空格,并且? 已经被丢弃,因为它是字段分隔符。

我假设在每一行中总是有一个 ) 后跟空格和一个 ?。如果这不成立,我的解决方案可能需要更改。

(在 GNU Awk 5.0.1 中测试)

【讨论】:

  • 您的回答中没有特定于 GNU awk 的内容,它适用于任何 awk,因为所有 awk 都将多字符 FS 视为正则表达式。
猜你喜欢
  • 2021-12-06
  • 1970-01-01
  • 2023-02-16
  • 2021-12-25
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-04-17
  • 1970-01-01
相关资源
最近更新 更多