【问题标题】:Extract text between 2 similar or different strings separately in shell script在shell脚本中分别提取2个相似或不同字符串之间的文本
【发布时间】:2022-01-12 08:54:59
【问题描述】:

我想分别提取每个 ### 之间的文本以与不同的文件进行比较。需要提取所有 docker 图像的所有 CVE 数字,以便与之前的报告进行比较。文件如下所示。这是一个 sn-p,它有 100 多个这样的行。需要通过 Shell 脚本执行此操作。请帮忙。

### Vulnerabilities found in docker image alarm-integrator:22.0.0-150
| CVE  | X-ray Severity | Anchore Severity | Trivy Severity | TR   |
| :--- | :------------: | :--------------: | :------------: | :--- |
|[CVE-2020-29361](#221fbde4e2e4f3dd920622768262ee64c52d1e1384da790c4ba997ce4383925e)|||Important|
|[CVE-2021-35515](#898e82a9a616cf44385ca288fc73518c0a6a20c5e0aae74ed8cf4db9e36f25ce)|||High|

### Vulnerabilities found in docker image br-agent:22.0.0-154
| CVE  | X-ray Severity | Anchore Severity | Trivy Severity | TR   |
| :--- | :------------: | :--------------: | :------------: | :--- |
|[CVE-2020-29361](#221fbde4e2e4f3dd920622768262ee64c52d1e1384da790c4ba997ce4383925e)|||Important|
|[CVE-2021-23214](#75eaa96ec256afa7bc6bc3445bab2e7c5a5750678b7cda792e3c690667eacd98)|||Important|

我尝试过类似grep -oP '(?<=\"##\").*?(?=\"##\")' 的方法,但它不起作用。

预期输出:

For alarm-integrator
CVE-2020-29361
CVE-2021-35515

For br-agent
CVE-2020-29361
CVE-2021-23214

【问题讨论】:

  • 能否请您更清楚地在您的问题中发布预期输出的示例,以使问题更清楚。谢谢
  • 嗨拉文德,感谢您的回复。我已经更新了预期的输出。请检查。

标签: shell awk sed grep script


【解决方案1】:

使用您展示的示例,请尝试关注awk 代码。

awk '
/^##/ && match($0,/docker image[[:space:]]+[^:]*/){
  split(substr($0,RSTART,RLENGTH),arr1)
  print "for "arr1[3]
  next
}
match($0,/^\|\[[^]]*/){
  print substr($0,RSTART+2,RLENGTH-2)
}
'  Input_file

说明:为上述awk代码添加详细说明。

awk '                                   ##Starting awk program from here.
/^##/ && match($0,/docker image[[:space:]]+[^:]*/){  ##Checking condition if line starts from ## AND using match function to match regex docker image[[:space:]]+[^:]* to get needed value.
  split(substr($0,RSTART,RLENGTH),arr1) ##Splitting matched part in above match function into arr1 array with default delimiter of space here.
  print "for "arr1[3]                   ##Printing string for space arr1 3rd element here
  next                                  ##next will skip all further statements from here.
}
match($0,/^\|\[[^]]*/){                 ##using match function to match starting |[ till first occurrence of ] here.
  print substr($0,RSTART+2,RLENGTH-2)   ##printing matched sub string from above regex.
}
'  Input_file                           ##mentioning Input_file name here.

【讨论】:

    【解决方案2】:

    使用 GNU awk(我假设您已经拥有或可以获得,因为您使用的是 GNU grep)作为第三个参数来匹配():

    $ cat tst.awk
    match($0,/^###.* ([^:]+):.*/,a) { print "For", a[1] }
    match($0,/\[([^]]+)/,a)         { print a[1] }
    !NF
    

    $ awk -f tst.awk file
    For alarm-integrator
    CVE-2020-29361
    CVE-2021-35515
    
    For br-agent
    CVE-2020-29361
    CVE-2021-23214
    

    【讨论】:

      【解决方案3】:

      awk 你可以这样做:

      awk -v FS=' |[[]|[]]' '/^[#]+/{sub(/:.*$/,"");print "For " $NF} /^\|\[/{print  $2} /^$/ {print ""}' file
      For alarm-integrator
      CVE-2020-29361
      CVE-2021-35515
      
      For br-agent
      CVE-2020-29361
      CVE-2021-23214
      
      • 我们将字段分隔符FS 配置为 |[[]|[]]:空格或[ 字符或] 字符。
      • 第一个条件操作用于获取For alarm-integratorFor br-agent
      • all CVE numbers 的第二个条件操作
      • 最后我们添加空行。

      更具可读性:

      awk -v FS=' |[[]|[]]' '
      /^[#]+/{sub(/:.*$/,"");print "For " $NF}
      /^\|\[/{print  $2}
      /^$/ {print ""}
      ' file
      For alarm-integrator
      CVE-2020-29361
      CVE-2021-35515
      
      For br-agent
      CVE-2020-29361
      CVE-2021-23214
      

      【讨论】:

        猜你喜欢
        • 2014-11-05
        • 2015-05-14
        • 2018-05-03
        • 2016-10-27
        • 2016-08-20
        • 1970-01-01
        • 2020-03-10
        • 1970-01-01
        • 2019-03-03
        相关资源
        最近更新 更多