【问题标题】:Exclude e-mails which domain name match with the global one排除域名与全局匹配的邮件
【发布时间】:2012-10-18 15:50:08
【问题描述】:

全局域在“*@”选项中,当电子邮件与这些全局域之一匹配时,我需要将它们从列表中排除。

例子:

WF,*@stackoverflow.com
WF,*@superuser.com
WF,*@stackexchange.com
WF,test@superuser.com
WF,test@stackapps.com
WF,test@stackexchange.com

输出:

WF,*@stackoverflow.com
WF,*@superuser.com
WF,*@stackexchange.com
WF,test@stackapps.com

【问题讨论】:

  • 全球域是否总是在电子邮件地址之前?
  • 在这种情况下是的,但在未来不会。

标签: sed awk grep cut tr


【解决方案1】:

你可以这样做:

grep -o "\*@.*" file.txt | sed -e 's/^/[^*]/' > global.txt
grep -vf global.txt file.txt

这将首先提取全局电子邮件,并在它们前面加上[^*],将结果保存到global.txt。然后将该文件用作 grep 的输入,其中每一行都被视为[^*]*@global.domain.com 形式的正则表达式。 -v 选项告诉 grep 只打印与该模式不匹配的行。

使用 sed 进行就地编辑的另一个类似选项是:

grep -o "\*@.*" file.txt | sed -e 's/^.*$/\/[^*]&\/d/' > global.sed
sed -i -f global.sed file.txt

【讨论】:

    【解决方案2】:
    $ awk -F, 'NR==FNR && /\*@/{a[substr($2,3)]=1;print;next}NR!=FNR && $2 !~ /^\*/{x=$2;sub(/.*@/,"",x); if (!(x in a))print;}' OFS=, file file
    WF,*@stackoverflow.com
    WF,*@superuser.com
    WF,*@stackexchange.com
    WF,test@stackapps.com
    

    【讨论】:

    • 对我的解释方式非常有帮助和简单。
    【解决方案3】:

    你在同一个文件中有两种类型的数据,所以最简单的处理方法是先分割:

    <infile tee >(grep '\*@' > global) >(grep -v '\*@' > addr) > /dev/null
    

    然后使用global 删除addr 中的信息:

    grep -vf <(cut -d@ -f2 global) addr
    

    把它放在一起:

    <infile tee >(grep '\*@' > global) >(grep -v '\*@' > addr) > /dev/null
    cat global <(grep -vf <(cut -d@ -f2 global) addr) > outfile
    

    outfile的内容:

    WF,*@stackoverflow.com
    WF,*@superuser.com
    WF,*@stackexchange.com
    WF,test@stackapps.com
    

    使用rm global addr清理临时文件。

    【讨论】:

    • 感谢您的精彩解释。
    【解决方案4】:

    这是使用GNU awk 的一种方式。运行如下:

    awk -f script.awk file.txt{,}
    

    script.awk的内容:

    BEGIN {
        FS=","
    }
    
    FNR==NR {
        if (substr($NF,1,1) == "*") {
            array[substr($NF,2)]++
        }
        next
    }
    
    substr($NF,1,1) == "*" || !(substr($NF,index($NF,"@")) in array)
    

    结果:

    WF,*@stackoverflow.com
    WF,*@superuser.com
    WF,*@stackexchange.com
    WF,test@stackapps.com
    

    或者,这里是单行:

    awk -F, 'FNR==NR { if (substr($NF,1,1) == "*") array[substr($NF,2)]++; next } substr($NF,1,1) == "*" || !(substr($NF,index($NF,"@")) in array)' file.txt{,}
    

    【讨论】:

      【解决方案5】:

      通过一次文件并允许将全局域与地址混合:

      $ cat file
      WF,*@stackoverflow.com
      WF,test@superuser.com
      WF,*@superuser.com
      WF,test@stackapps.com
      WF,test@stackexchange.com
      WF,*@stackexchange.com
      WF,foo@stackapps.com
      $
      $ awk -F'[,@]' '
         $2=="*" { glbl[$3]; print; next }
         { addrs[$3] = addrs[$3] $0 ORS }
         END {
            for (dom in addrs)
               if (!(dom in glbl))
                  printf "%s",addrs[dom]
         }
      ' file
      WF,*@stackoverflow.com
      WF,*@superuser.com
      WF,*@stackexchange.com
      WF,test@stackapps.com
      WF,foo@stackapps.com
      

      或者如果您不介意 2-pass 方法:

      $ awk -F'[,@]' '(NR==FNR && $2=="*" && !glbl[$3]++) || (NR!=FNR && !($3 in glbl))' file file
      WF,*@stackoverflow.com
      WF,*@superuser.com
      WF,*@stackexchange.com
      WF,test@stackapps.com
      WF,foo@stackapps.com
      

      我知道第二个有点神秘,但它很容易翻译为不使用默认操作和 awk 习语的一个很好的练习:-)。

      【讨论】:

        【解决方案6】:

        这可能对你有用(GNU sed):

        sed '/.*\*\(@.*\)/!d;s||/[^*]\1/d|' file | sed -f - file
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 2016-03-20
          • 1970-01-01
          • 2014-05-18
          • 2011-07-31
          • 1970-01-01
          • 2019-12-15
          • 2012-12-28
          相关资源
          最近更新 更多