【问题标题】:Formatting text outputs in unix在 unix 中格式化文本输出
【发布时间】:2017-11-01 02:39:40
【问题描述】:

您好,我这里有一个清单:

list_1.txt

Alpha

Bravo

Charlie

以及目录中具有以下文件名和内容的文件:

Alpha_123.log

This is a sample line in the file
error_log "This is error1 in file"
This is another sample line in the file
This is another sample line in the file
This is another sample line in the file
error_log "This is error2 in file"
This is another sample line in the file
This is another sample line in the file
error_log "This is error3 in file"
This is another sample line in the file
This is another sample line in the file

Alpha_123.sh

This is a sample line in the file
This is another sample line in the file
This is another sample line in the file
error_log "This is errorA in file"
This is another sample line in the file
This is another sample line in the file
This is another sample line in the file
This is another sample line in the file
error_log "This is errorB in file"
This is another sample line in the file
This is another sample line in the file
error_log "This is errorC in file"
This is another sample line in the file

Bravo.log

Charlie.log

Bravo.log 和 Charlie.log 的内容与 Alpha.log 类似

我想要这样的输出:

Alpha|"This is error1 in file"

Alpha|"This is error2 in file"

Alpha|"This is error3 in file"

Alpha|"This is errorA in file"

Alpha|"This is errorB in file"

Alpha|"This is errorC in file"

非常感谢任何输入。谢谢!

所以基本上,我想首先在list_1.txt中找到名称包含字符串模式的文件,然后找到错误消息并使用|输出

【问题讨论】:

  • 您可以从格式化问题开始 - stackoverflow.com/editing-help 并添加您尝试解决此问题的内容..
  • 您的问题不清楚。更新您的描述
  • 你想做什么
  • 基本上,我想在list_1.txt中找到带有字符串模式的日志文件。然后在这些文件中搜索包含“error_log”的行,并输出包含 list_1.txt 中的字符串模式和包含用管道分隔的“error_log”的行。
  • 我有这个但是这个:cat list_1.txt | awk '{打印 $1}' |同时读取 x ;回声 $x"|" ls log_directory | grep $x | xargs -i grep error_log log_directory {} | egrep -v ; done 但是输出是这样的: Alpha|error_log "This is error1 in file" error_log "This is error2 in file" error_log "This is error3 in file" 我想要的是这样的:Alpha|"This is error1 in file" Alpha |"这是文件中的错误 2"Alpha|"这是文件中的错误 3"

标签: unix awk sed grep text-processing


【解决方案1】:

如果我理解正确,应该这样做:

 awk -vOFS=\| 'FNR==1{file=FILENAME;sub("[.]log$","",file)}{sub("^error_log ","");print file,$0}' *.log

解释:

  • -vOFS=\| 将输出字段分隔符设置为|。 (\ 需要从 shell 中转义 |(这会将其视为管道)。您可以改用 -vOFS='|'。)
  • FNR==1{...} 确保此代码在每个输入文件中仅运行一次:FNR 是 awk 从当前文件读取的记录(即行)数.因此,在处理每个文件的第一行时,这仅等于 1。
  • file=FILENAME 只是将当前处理的输入文件的文件名存储在一个变量中以供以后编辑。
  • sub("[.]log$","",file) 删除 .log([...] 将点 (.) 转义为正则表达式中的任何字符。您可以使用 \\.取而代之。)从文件名的末尾(这就是$ 所代表的)开始。
  • {...} 为每个输入文件的每个记录/行运行代码。
  • sub("^error_log ","") 从输入的每一行(“record”)的开头(这就是^ 所代表的意思)删除"error_log "(注意尾随空格!)。
  • print file,$0 打印每个 record 的剩余部分(即行),以相应的文件名为前缀。请注意,逗号 (,) 将替换为前面指定的 输出字段分隔符。您可以使用 print file "|" $0 而不是指定 OFS。
  • *.log 将使当前目录中以.log 结尾的每个文件成为awk 命令的输入文件。您可以改为明确指定 Alpha.log Bravo.log Charly.log。

这是使用您的 list.txt 构造文件名的替代方法:

awk -vOFS=\| '{file=$0;logfile=file ".log";while(getline < logfile){sub("^error_log ","");print file,$0}}' list.txt

解释:

  • file=$0 将来自list.txt 的当前行(record)保存在一个变量中。
  • logfile=file ".log" 将.log 附加到它以获取相应的日志文件名。
  • while(getline &lt; logfile){...} 将为当前日志文件中的每一行/记录运行代码。

上面的例子应该很清楚了。

【讨论】:

    【解决方案2】:

    awk 来救援!

    awk '{gsub(/^error_log /,FILENAME"|")}1' $(awk '{print $0".log"}' list_1.txt)
    

    更新

    根据更新的信息,我认为这就是您要查找的内容。

    awk '/^error_log/ {split(FILENAME,f,"_"); 
                       gsub(/^error_log /,f[1]"|")}' $(awk '{print $0"_*"}' list_1.txt)
    

    【讨论】:

    • 嗨@karaka,感谢您的意见。不幸的是,这并没有考虑到使用 list_1.txt 中的字符串查找日志文件
    • 你试过了吗?它按照你的要求做。
    • 嗨 Karaka,抱歉我错过了我的问题的一些细节,请查看更新后的问题。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-09-22
    • 2020-02-26
    相关资源
    最近更新 更多