【问题标题】:how to check for the presence of two words in a file using GREP如何使用 GREP 检查文件中是否存在两个单词
【发布时间】:2016-09-16 01:32:59
【问题描述】:

我有两个文件 A.txt 和 B.txt 分别包含两个列表,如下所示。

文件 A.txt

hello
hi 
ko

文件 B.txt

fine
No 
And how
why

现在我想检查在另一个文件 C.txt 的一行中是否存在任何这些单词(来自 A.txt 和 B.txt)。

我正在使用 grep 命令

grep -iof A.txt C.txt| grep B.txt

C.txt 包含包含 A.txt 和 B.txt 中单词的句子

Hello I am fine
I am not fine
why ko is and how?

不显示任何输出

所以,现在我想如果 A.txt 和 B.txt 中的任何单词同时出现在一个句子中,它应该将输出显示为

Hello fine
why ko and how 

如果两个文件同时出现在 C.txt 中,则只打印匹配的单词,而不是打印 C.txt 中的整行

【问题讨论】:

  • 您说存在其中一个这些词。我不确定我是否理解 either 在这种情况下。你的意思是存在任何这些词,不是吗? IE。您想知道 C.txt 中是否有一行包含来自 A.txt 或来自 B.txt 的单词。对吗?
  • 我刚刚编辑了这个问题。我的意思是 A.txt 和 B.txt 中的任何这些词在 C.txt 中同时出现
  • 在这种情况下,我不太明白您 simultaneously 是什么意思,但是既然您已经接受了答案,我想这就是您要找的。 ;-)

标签: regex shell scripting grep


【解决方案1】:

你可能想说:

$ grep -if B <(grep -if A C)
Hello I am fine
why ko is and how?

这使用-f 来提供表达式。它可以是一个文件...或您使用process substitution &lt;( ... ) 即时创建的文件。

首先,grep -if A C 匹配C 中所有在A 中的单词:

$ grep -if A C
Hello I am fine        # "Hello" highlighted
why ko is and how?     # "ko" highlighted

然后,将其输出与B中的内容进行比较。

$ grep -if B <(grep -if A C)
Hello I am fine        # "fine" highlighted
why ko is and how?     # "and how" highlighted

根据您的需要,您可能需要添加-F、-w 和-i。

来自man grep:

   -f FILE, --file=FILE
          Obtain  patterns  from  FILE,  one  per  line.   The  empty file
          contains zero patterns, and therefore matches nothing.   (-f  is
          specified by POSIX.)

   -F, --fixed-strings
          Interpret PATTERN as a  list  of  fixed  strings,  separated  by
          newlines,  any  of  which is to be matched.  (-F is specified by
          POSIX.)

   -i, --ignore-case
          Ignore  case  distinctions  in  both  the  PATTERN and the input
          files.  (-i is specified by POSIX.)

   -w, --word-regexp
          Select  only  those  lines  containing  matches  that form whole
          words.  The test is that the matching substring must  either  be
          at  the  beginning  of  the  line,  or  preceded  by  a non-word
          constituent character.  Similarly, it must be either at the  end
          of  the  line  or  followed by a non-word constituent character.
          Word-constituent  characters  are  letters,  digits,   and   the
          underscore.

【讨论】:

  • 我在哪里提到这两个文件?
  • 我认为应该是 - 为了安全起见 - fgrep,而不是 grep - 以防万一模式文件中包含句点或 $ 等字符。另一个问题是 OP 想要匹配 words,而不是字符串。 Grepping for,即“hi”,也将匹配例如“hint”,这可能不是 OP 想要的。无论如何,我认为在我们提出解决方案之前应该更准确地说明问题。
  • 我刚刚编辑过,主要目的是找出 A.txt 和 B.txt 中的任何单词在 C.txt 中同时出现
  • @Doej 用$ grep -f B &lt;(grep -if A C)查看我的更新。
  • @user1934428 好点。我添加(或建议)-F、-i、-w。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2013-02-09
  • 1970-01-01
  • 2012-11-09
  • 1970-01-01
  • 2014-04-06
  • 1970-01-01
  • 2018-04-17
相关资源
最近更新 更多