【问题标题】:sed regex: group repeat option?sed 正则表达式:组重复选项?
【发布时间】:2019-02-07 15:39:07
【问题描述】:

我有一个包含几行的文本输入。每个组由一个空行 (\n\n) 分隔。 我正在使用 sed 进行处理,但我对替代方案持开放态度。

我正在使用这个结构来一次处理所有的行:

# if the first line copy the pattern to the hold buffer
1h
# if not the first line then append the pattern to the hold buffer
1!H
# if the last line then ...
$ {
  # copy from the hold to the pattern buffer
  g

  ... here are my regex lines.

  # print
  p
}

我对每个组的目标输出是每一行,但第一行以第一行的内容为前缀,以空格分隔。

由于我当前的输入只有 2、3 和 6 行组,因此我“硬编码”了它 像这样:

2 行: s/\n\n\([^\n]\+\)\n\([^\n]\+\)\n\n/\n\n\1 \2\n\n/g

3 行: s/\n\n\([^\n]\+\)\n\([^\n]\+\)\n\([^\n]\+\)\n\n/\n\n\1 \2\n\n\1 \3\n\n/g

6 行: s/\n\n\([^\n]\+\)\n\([^\n]\+\)\n\([^\n]\+\)\n\([^\n]\+\)\n\([^\n]\+\)\n\([^\n]\+\)\n\n/\n\n\1 \2\n\n\1 \3\n\n\1 \4\n\n\1 \5\n\n\1 \6\n\n/g

(我有两次这些正则表达式行,因为可能需要一组的结尾 \n\n 并且不能用于匹配下一组的开头)

我正在寻找一种适用于从 2 到 n 行的任何大小的组的通用方法。有人对此有任何想法吗?

更新:因为@Benjamin W. 要求样本输入/输出:

我在这里要解决的真正问题是为温度记录守护程序动态生成一个 csv 标题行,该守护程序的数据来源于sensors -u。 (因为我的笔记本电脑重启时输出的顺序似乎发生了变化)

使用 sed 很容易从原始程序输出到这个:

jc42-i2c-0-1a SMBus I801 adapter at f040
temp1

asus-isa-0000 ISA adapter
cpu_fan
temp1

acpitz-acpi-0 ACPI interface
temp1

jc42-i2c-0-18 SMBus I801 adapter at f040
temp1

coretemp-isa-0000 ISA adapter
Package id 0
Core 0
Core 1
Core 2
Core 3

我上面提到的 3 个 sed 正则表达式替换行允许我将其转换为:

jc42-i2c-0-1a SMBus I801 adapter at f040 temp1
asus-isa-0000 ISA adapter cpu_fan
asus-isa-0000 ISA adapter temp1
acpitz-acpi-0 ACPI interface temp1
jc42-i2c-0-18 SMBus I801 adapter at f040 temp1
coretemp-isa-0000 ISA adapter Package id 0
coretemp-isa-0000 ISA adapter Core 0
coretemp-isa-0000 ISA adapter Core 1
coretemp-isa-0000 ISA adapter Core 2
coretemp-isa-0000 ISA adapter Core 3

但这当然只适用于具有 1、2 或 5 个值的适配器的机器。

2019-02-11 更新:

所以在我得到两个建议通用解决方案的答案后,我再次查看了这个问题并简化了我的整个温度记录脚本:

echo -n "timestamp"
sensors -u | # -u gives Raw output, suitable for easier post-processing
grep --invert-match '  ' | # remove all lines containing values, leaving only headers
sed -n 'H; ${x; s/\nAdapter: / /g; p}' | # join headers spanning two lines together. For syntax see: https://unix.stackexchange.com/questions/163428/replace-a-string-containing-newline-characters & http://www.grymoire.com/Unix/Sed.html#uh-55
sed 'N;/\n$/d;s/\(.*\)\n\(.*\):/\1 \2\n\1/;P;$d;D' | # join the headers header with each sub-header, see: https://stackoverflow.com/questions/54576948/sed-regex-group-repeat-option
tr '\n' ';' | sed 's/.$//' # join finished headers together in a single line sepearted by ; & remove the trailing ;
echo ""

while true
do
    ts=`date +"%Y-%m-%d %H:%M:%S"`
    echo -n "$ts;"
    sensors -u | grep --invert-match '_max\|_crit\|_min' | # remove min max crit values which represent config, not state.
    grep '\.' | # remove all non value lines left (headers & empty lines seperating blocks
    sed 's/  .*: //g' | # remove value names, leaving only the values themselfs
    sed 's/\.000//g' | # remove empty decimals
    tr '\n' ';' | sed 's/.$//' # join finished values together in a single line sepearted by ; & remove the trailing ;
    sleep 1
    echo ""
done

【问题讨论】:

  • awk RS="\n\n"?
  • 给我syntax error.. 抱歉,我没有使用 awk 的经验,你必须给我更多。
  • 这只是您的搜索词。我们无法确切知道您要做什么,但 awk -v RS="\n\n" '{ something with each block }' 听起来很像您正在寻找的东西。
  • @tripleee 对于空行分隔的记录,您可以使用RS = ""(参见manual)。
  • 你能添加示例输入和输出吗?

标签: regex bash sed


【解决方案1】:

这可能对你有用(GNU sed):

sed 'N;/\n$/d;s/\(.*\)\n\(.*\)/\1 \2\n\1/;P;$d;D' file

将下一行添加到当前行。

如果附加的行是空的,即\n$ 表示空行,则完全删除模式空间并恢复,就像没有行被消耗一样。

否则,模式空间中的两行都是非空的,所以将两行转换为一个,然后将第一行附加到结果中。

打印模式空间中的第一行。

如果是文件的最后一行,删除模式空间。

删除模式空间中的第一行。

重复。

注意D 删除模式空间中的第一行,如果模式空间不为空,则不会隐式将模式空间替换为下一行。

【讨论】:

  • 感谢您的解决方案!非常紧凑并且使用了我要求的工具,尽管 awk 解决方案在我看来更具可读性,我对 sed 有一些经验,而对 awk 几乎没有经验。但一个问题是,您的两个解决方案都在完成最后一个块后再次打印最后一个块的第一行。
  • @kaefert 哎呀!忘记删除最后一行了。
【解决方案2】:

这可以作为 awk 解决方案:

awk 'BEGIN {RS="\n\n"; FS="\n"} {for (i = 2; i <= NF; i++) print $1,$i}' file
  • 将“\n\n”定义为记录分隔符 (RS)
  • 将“\n”定义为字段分隔符 (FS)
  • 对于每条记录中从第二个到最后一个 (NF) 的每个字段:打印第一个字段 ($1) 和当前字段 ($i),由 OFS 连接,由 "," 触发

【讨论】:

  • 感谢您的解决方案并解释发生了什么!伟大的!在我看来有点奇怪的是循环中的 +i++ 部分。 (你能解释一下它的作用吗?)另外(就像@potong 的解决方案一样)它在完成最后一个块后再次打印最后一个块的第一行。
  • 我试过这个,但没有解决问题:awk 'BEGIN {RS="\n\n"; FS="\n"; OFS="\n"} {if (NF > 1) {for (i = 2; i
  • 嗨,kaefert,i 前面的附加 + 是一个错字,显然没有效果。我删除了它。
  • 我认为最后一个块中的额外打印是因为在输入文件中文件末尾之前的最后一行末尾有一个换行符。代码行为正确,因为您有一个空字段作为最后一个元素。查看here 以使用常用编辑器调整您的输入。
  • OFS 在我编写的第一个版本中毫无用处。我把它改成了有用的版本。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2012-11-17
  • 2018-11-11
  • 1970-01-01
  • 1970-01-01
  • 2010-12-31
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多