【问题标题】:Filtering with awk based on two criteria根据两个条件使用 awk 进行过滤
【发布时间】:2017-08-25 13:33:36
【问题描述】:

我之前也有类似的问题,但这次我需要一些更复杂的东西:

在一个看起来像这样的 txt 文件中:

147 186741 2S74M -162
83 647172 1S75M -221
163 584665 74M2S 271
99 658416 5S65M6S -272
163 718735 60M16S 243

我希望 awk 查看第三列,当它在第二或第三位置遇到字符“S”时,它会查看第一列,当它遇到“147”或“83”时,它会丢弃那条线。其余结果被传递给第二个 awk,它再次查看第 3 行,当它在最后遇到字符“S”时,它会查看第一列,如果找到“99”或“163”它会丢弃这些行。然后它打印不符合这些过滤器的其余行。

我尝试了一些类似的方法,但得到了空白文件:

awk -Ft '{if ($3 ~ /S$/  && $1 ~ /99|163/)} {next}' | awk -Ft '{if ($3 ~ /^..?S/ && $1 ~ /147|83/)} {next} $6 ~ /S/ {print}' input.txt > output.txt 

【问题讨论】:

  • when it encounters 是什么意思?如果 $1 正好是123147456 那么147 是否“遇到”了?请使用“完全匹配”或“包含”等术语(以及部分匹配和完全匹配以及字符串或数字或正则表达式比较),并显示与您的标准不匹配的示例,尤其是下雨天/边缘情况。另外,当您的输入文件中没有ts 并且如果有它真的会破坏您的脚本时,您为什么要设置-Ft。最后 - 你的第一个 awk 语句没有打印任何输出,也没有读取任何输入文件,所以你当然会得到一个空白的输出文件。

标签: bash awk


【解决方案1】:

由于您没有显示您正在使用的 Input_file,所以我根据您显示的 Input_file 举了我的例子,假设以下是您的 Input_file。

cat Input_file
147 186741 2S74M -162
83 647172 1S75M -221
163 584665 74M2S 271
99 658416 5S65M6S -272
163 718735 60M16S 243
147 186741 2K74M -162
83 647172 1K75M -221
163 584665 74M2K 271
99 658416 5S65M6K -272
163 718735 60M16S 243

下面是我的代码:

awk '(($1==147 || $1==83) && (substr($3,2,1)=="S" || substr($3,3,1)=="S")) || (substr($3,length($3))=="S" && ($1==99 || $1==163)){next} 1'  Input_file

现在,当我在 awk 上运行时,我得到的这些值(我添加这些值只是为了检查我的代码是否正常工作)如下。

awk '(($1==147 || $1==83) && (substr($3,2,1)=="S" || substr($3,3,1)=="S")) || (substr($3,length($3))=="S" && ($1==99 || $1==163)){next} 1' Input_file
147 186741 2K74M -162
83 647172 1K75M -221
163 584665 74M2K 271
99 658416 5S65M6K -272

所以你可以看到所有那些不符合你提供的条件的行正在打印,给我一些时间也会在这里添加解释。

编辑:在这里也添加对上述代码的解释,请不要运行它,因为我已将其划分为不同的部分,仅供 OP 理解。

awk '(($1==147 || $1==83)\  ##First condition which re-presents your first awk starts here. checking conditions where $1 value is either 147 OR $1 value is 83
&& \                        ##  AND
(substr($3,2,1)=="S" \      ##substring of 3rd column is EQUAL to letter S
|| \                        ##  OR
substr($3,3,1)=="S"))\      ##substring of 3rd column is EQUAL to letter S
|| \                        ##OR(means either that first aw condition should be TRUE or this following one), the second major condition for which you used second awk I clubbed both the awks into 2 major conditions here.
(substr($3,length($3))=="S"\##checking if substring of column 3s last letter is EQUAL to S here.
&& \                        ##  AND
($1==99 || $1==163)){       ##$1 value is either 99 or 163. So if either of above 2 major conditions are TRUE then perform following statements.
next                        ##next, it is awk keyword which will skip all further statements of line now, without doing any action.
}
1                           ##awk works on method of condition and then action, so here I am making condition as TRUE by mentioning as 1 and NO action is mentioned so be default print action will happen which will print current line.
' Input_file                  ##mentioning Input_file here.

【讨论】:

  • 这和以前的解决方案一样好用(我仍然需要做一个额外的过滤步骤,但我忘了提到,但它应该可以工作)。我喜欢您使用“substr”的方式,并感谢您的额外解释,因为我刚刚从您的帖子中学到了很多东西(作为编码新手,拥有这些真的很有帮助)!非常感谢,非常感谢。
【解决方案2】:

对于初学者来说,$6 可能是一个错字。

现在让我们尝试分步执行此操作。第 1 步:

awk '$1 ~ /147|83/ && $3 ~ /^..?S/ {next;} {print;}' test.txt

给我们留下:

163 584665 74M2S 271
99 658416 5S65M6S -272
163 718735 60M16S 243

如果将这些行放在文件 test2.txt 中,则应用:

awk '($1 ~ /99|163/ && $3 ~ /S$/) {next;} {print;}' test2.txt

没有有效的行,因为所有第 3 列的末尾都有一个“S”,并且以 99 或 163 开头。

【讨论】:

  • ...您可以一次性完成这两个语句。比如 awk '($1 ~ /147|83/ && $3 ~ /^..?S/) || ($1 ~ /99|163/ && $3 ~ /S$/) {next;} {print;}' test.txt
  • 是的,这两种方法都很好,我更喜欢单行语句。由于我忘记提及第三个标准,我将不得不对其进行一些修改,在该标准中我将删除所有根本没有“S”的读取,但我想我可以解决这个问题。非常感谢!!!
  • NP。如果您遇到困难,请告诉我们。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-05-25
  • 2023-03-11
  • 1970-01-01
  • 2011-03-24
  • 1970-01-01
  • 2019-06-14
相关资源
最近更新 更多