【问题标题】:How to find rows containing integers given in a file?如何查找包含文件中给定整数的行?
【发布时间】:2012-08-16 22:14:19
【问题描述】:

我有一个文件dict,每行包含一个整数

123
456

我想在文件file 中找到恰好包含dict 中的整数的行。

如果我使用

$ grep -w -f dict file

我得到错误的匹配,例如

12345  foo
23456  bar

这些都是错误的,因为 12345 != 123 和 23456 != 456。问题是-w 选项也将数字视为单词字符。 -x 选项也不起作用,因为file 中的行可以有其他文本。请问最好的方法是什么?如果该解决方案能够在大尺寸的dict 和file 上提供进度监控和良好的性能,那就太好了。

【问题讨论】:

  • 必须使用grep,还是愿意接受其他解决方案?
  • 并非如此。任何命令行工具都可以。
  • 你的grep 命令对我有用,没有你指出的误报。

标签: regex search awk grep pattern-matching


【解决方案1】:

将单词边界添加到dict中,如下所示:

\<123\>
\<456\>

-w 参数不是必需的。只需要:

grep -f 字典文件

【讨论】:

  • 这行得通。但是,当我尝试使用包含 20000 行的 file 时,速度非常慢。有什么提高性能的建议吗?另外,当它非常慢时,是否可以监视正在处理字典中的哪个条目,以便我对程序的进度有所了解?
  • 请将所有模式写入一行:\\|\
【解决方案2】:

使用awk的相当通用的方法:

awk 'FNR==NR { array[$1]++; next } { for (i=1; i<=NF; i++) if ($i in array) print $0 }' dict file

说明:

FNR==NR { }  ## FNR is number of records relative to the current input file. 
             ## NR is the total number of records.
             ## So this statement simply means `while we're reading the 1st file
             ## called dict; do ...`

array[$1]++; ## Add the first column ($1) to an array called `array`.
             ## I could use $0 (the whole line) here, but since you have said
             ## that there will only be one integer per line, I decided to use
             ## $1 (it strips leading and lagging whitespace; if any)

next         ## process the next line in `dict`

for (i=1; i<=NF; i++)  ## loop through each column in `file`

if ($i in array)       ## if one of these columns can be found in the array

print $0               ## print the whole line out

使用 bash 循环处理多个文件:

## This will process files; like file, file1, file2, file3 ...
## And create output files like, file.out, file1.out, file2.out, file3.out ...

for j in file*; do awk -v FILE=$j.out 'FNR==NR { array[$1]++; next } { for (i=1; i<=NF; i++) if ($i in array) print $0 > FILE }' dict $j; done

如果您有兴趣在多个文件上使用tee,您可能想尝试这样的事情:

for j in file*; do awk -v FILE=$j.out 'FNR==NR { array[$1]++; next } { for (i=1; i<=NF; i++) if ($i in array) { print $0 > FILE; print FILENAME, $0 } }' dict $j; done 2>&1 | tee output

这将显示正在处理的文件的名称和找到的匹配记录,并将“日志”写入名为 output 的文件。

【讨论】:

  • 你能解释一下这个程序是如何工作的吗?
  • 另外,awk 方法能否监控进度?
  • @user001:如果dict 非常大,您将无法知道array 中添加了多少数字。但是在读取file时,如果找到匹配的,它会立即打印出匹配的行。
  • 是的,我注意到它会立即打印出匹配的结果。这是一个很好的功能。顺便说一句,是否可以打印匹配的行并同时将其写入文件?
  • 另外,NR==FNR {} 用于判断当前行是否在第一个文件中对我来说有点奇怪。说,如果我想扩展程序以处理多个文件file1,file2 与单个dict 文件,这是行不通的。
【解决方案3】:

您可以使用 Python 脚本相当容易地做到这一点,例如:

import sys

numbers = set(open(sys.argv[1]).read().split("\n"))
with open(sys.argv[2]) as inf:
    for s in inf:
        if s.split()[0] in numbers:
            sys.stdout.write(s)

错误检查和恢复留给读者实现。

【讨论】:

  • 嗯,理想情况下,我想在 Bash 命令行上使用 GNU 实用程序来完成。如果有很多行代码,我可以做一个脚本。谢谢。
猜你喜欢
  • 2019-06-26
  • 2011-09-03
  • 2015-04-06
  • 2012-09-22
  • 1970-01-01
  • 2010-12-17
  • 1970-01-01
  • 2019-04-28
  • 1970-01-01
相关资源
最近更新 更多