【问题标题】:How do I print the highest/longest values in a file如何打印文件中的最高/最长值
【发布时间】:2022-11-13 11:12:54
【问题描述】:

我有一个 AV 日志文件,其中显示了每个扫描的进程的多个值:名称、路径、扫描的文件总数、扫描时间。该文件包含数百个这样的流程条目(下面的示例)和扫描的文件总数和扫描时间我想排序并打印最高(或最长)的值,以便确定哪些进程正在影响系统。我用 grep 尝试了各种方法,但似乎只得到一个按数字顺序运行的列表,当我真正想说的是进程 id:86,扫描时间(ns):12761174 是最高的,然后是进程 id 25,等等. 希望我的解释足够清楚。

Process id: 25
Name: wwww
Path: "/usr/libexec/wwww"
Total files scanned: 42
Scan time (ns): "62416"
Status: Active

Process id: 7
Name: xxxx
Path: "/usr/libexec/xxxx"
Total files scanned: 0
Scan time (ns): "0"
Status: Active

Process id: 86
Name: yyyy
Path: "/usr/libexec/yyyy"
Total files scanned: 2
Scan time (ns): "12761174"
Status: Active

我努力了:

grep -Eo | grep 'Scan time (ns)' '[0-9]+' file | sort

结果是:

file:Scan time (ns): "9391986"
file:Scan time (ns): "9532119"
file:Scan time (ns): "9730650"
file:Scan time (ns): "9743828"
file:Scan time (ns): "9793469"
file:Scan time (ns): "9911768"

我想要实现的是:

Process id 9, Scan time (ns): "34561"
Process id 86, Scan time (ns): "45630"
Process id 25, Scan time (ns): "1256822"
Process id 51, Scan time (ns): "52351290"
Process id 30, Scan time (ns): "90257651"
Process id 19, Scan time (ns): "178764794932"

【问题讨论】:

  • 请更新问题以显示您的代码生成的(错误)输出和(正确)预期输出,确保两组输出与提供的示例输入相对应
  • 第一个 grep 有什么意义?
  • 想删掉,多此一举
  • 你从哪里得到 34561?请添加您想要的输出,以准确地将样本输入添加到您的问题中。
  • 显然grep 'Scan time (ns)' '[0-9]+' file 不起作用,因为grep 默认只接收1 个模式,其余非选项参数是输入文件。如果您希望grep 找到多个模式,那么您需要使用-e:grep -e 'Scan time (ns)' -e '[0-9]+' file,或者使用正则表达式模式:grep -P 'Scan time \(ns\)|[0-9]+' file

标签: bash sorting awk sed grep


【解决方案1】:

这是另一种方法。它使用sed 和sort:

sed '/^Process id:/h; /^Scan time (ns):/!d; s/"//g; H; x; s/
/, /' file | sort -k7,7n

注意:我已经删除了扫描时间值周围的双引号(整数值周围的双引号对我来说意义不大)。

【讨论】:

  • 这很好用,或多或少是我想要的,所以谢谢
【解决方案2】:

使用您显示的示例,请尝试遵循awk 代码。用 GNU awk 编写和测试。

awk '
/^Process id: /{
  val=$NF
  next
}
/^Scan time (ns): "/{
  arr[val]=$NF
}
END{
  PROCINFO["sorted_in"]="@ind_num_asc"
  for(i in arr){
    print "Process id " i ", Scan time (ns): " arr[i]""
}
}
'  Input_file

【讨论】:

  • 这输出干净,但没有给我我正在寻找的订单
【解决方案3】:

使用perl 一次读取一个记录(使用“段落模式”,它使用一个空行作为记录分隔符),提取时间,并按其倒序排序:

$ perl -00 -lne 'm/Scan time (ns):s+"(d+)"/ && push @procs, [ $_, $1 ];
                 END { print $_->[0] for sort { $a->[1] < $b->[1] } @procs }' input.txt
Process id: 86
Name: yyyy
Path: "/usr/libexec/yyyy"
Total files scanned: 2
Scan time (ns): "12761174"
Status: Active

Process id: 25
Name: wwww
Path: "/usr/libexec/wwww"
Total files scanned: 42
Scan time (ns): "62416"
Status: Active

Process id: 7
Name: xxxx
Path: "/usr/libexec/xxxx"
Total files scanned: 0
Scan time (ns): "0"
Status: Active

【讨论】:

  • 这可以干净地执行,但只包含文件,我看不到重新排序
  • @onthenile 仔细看
  • 我在您的示例中看到了逻辑,顺序相反,但是在我的文件的本地副本中,它有数百个进程 ID,这些 ID 仍然以数字方式运行,所以我没有看到低 @ 的有序列表987654323@ 值到高值(或相反)
【解决方案4】:

awk 的RS(输入记录分隔符)和FS(输入字段分隔符)的组合在这种情况下很有用:

< inpt awk 'BEGIN { RS = ""; FS = "
" } { print $1 ", " $5 }' | sort -t " -k2n
  • 在开始处理任何东西之前,即在BEGIN中,我们设置
    • RS 到 "",表示记录由空行分隔,并且
    • FS到" ",意思是(在每条记录中)字段之间用换行符分隔;
  • 然后我们继续打印第 1 和第 5 个字段,中间用逗号隔开。
  • 最后,将流的每一行解释为"-分隔的字段列表,我们根据第二个字段(-k2)进行sortn。

【讨论】:

  • 我的结果是=====================================, Total files scanned: 0。 “inpt”是文件名吗?我是这么认为的。
  • @onthenile,是的,inpt 是文件名。如果删除 | sort ... 会发生什么?你是什​​么系统?
  • 移除该管道后,结果没有变化。系统是 macOS,但使用 RHEL 是相同的输出。
  • @onthenile,echo -e '1 2 3 4 ' | awk 'BEGIN { RS = ""; ORS = "x" } 1 { print $1 }' 和 echo -e '1,2' | awk 'BEGIN { FS = ","; OFS = ";" } { print $1, $2 }' 打印什么?
【解决方案5】:

只有paste:

$ cat file | paste - - - - - - -
Process id: 25  Name: wwww  Path: "/usr/libexec/wwww"   Total files scanned: 42 Scan time (ns): "62416" Status: Active  
Process id: 7   Name: xxxx  Path: "/usr/libexec/xxxx"   Total files scanned: 0  Scan time (ns): "0" Status: Active  
Process id: 86  Name: yyyy  Path: "/usr/libexec/yyyy"   Total files scanned: 2  Scan time (ns): "12761174"  Status: Active  

如果我们添加一些格式:

$ cat file 
    | paste - - - - - - - 
    | awk '{
        printf("process id: %s, scan time (ns): %s
", $3, $15);
    }'
process id: 25, scan time (ns): "62416"
process id: 7, scan time (ns): "0"
process id: 86, scan time (ns): "12761174"

那是 7 个破折号 (-),因为您的每条记录都是 7 行(包括空行)。

破解说明:

paste 将所有输入文件的第一行连接成一行,然后连接第二行,依此类推。
因此,对于每个输入文件,它依次读取一行并将其添加到当前输出行。
我们已经将标准输入作为输入,7 次。但标准输入是一个单一的流。

所以paste会做:

  • 从输入 1 读取第 1 行(标准输入的第 1 行)
  • 从输入 2 读取第 1 行(标准输入的第 2 行)
  • ...
  • 从输入 7 读取第 1 行(标准输入的第 7 行)
  • 将这些行(stdin 的第 1-7 行)连接为 stdout 的第 1 行
  • 从输入 1(标准输入的第 8 行)读取第 2 行
  • ...
  • 从输入 7 读取第 2 行(标准输入的第 14 行)
  • 将这些行(标准输入的第 8-14 行)连接为标准输出的第 2 行
  • ...

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-06-12
    • 1970-01-01
    • 2017-02-08
    • 1970-01-01
    • 2021-08-17
    相关资源
    最近更新 更多