【问题标题】:awk Merge two files based on common field and print similarities and differencesawk 根据共同字段合并两个文件并打印异同
【发布时间】:2010-12-19 07:25:37
【问题描述】:

我有两个文件我想合并到第三个文件中,但我需要查看它们何时共享一个公共字段以及它们的不同之处。由于其他字段存在细微差异,我无法使用差异工具,我想这可以用 awk 完成。

文件 1:

aWonderfulMachine             1   mlqsjflk          
AnotherWonderfulMachine     2   mlksjf          
YetAnother WonderfulMachine 3   sdg         
TrashWeWon'tBuy             4   jhfgjh          
MoreTrash                     5   qsfqf         
MiscelleneousStuff           6  qfsdf           
MoreMiscelleneousStuff       7  qsfwsf

文件2:

aWonderfulMachine             22    dfhdhg          
aWonderfulMachine             23    dfhh            
aWonderfulMachine             24    qdgfqf          
AnotherWonderfulMachine     25    qsfsq         
AnotherWonderfulMachine     26    qfwdsf            
MoreDifferentStuff           27    qsfsdf           
StrangeStuffBought           28    qsfsdf

期望的输出:

aWonderfulMachine   1   mlqsjflk    aWonderfulMachine   22  dfhdhg
                                     aWonderfulMachine  23  dfhdhg
                                     aWonderfulMachine  24  dfhh
AnotherWonderfulMachine 2   mlksjf  AnotherWonderfulMachine 25  qfwdsf
                                       AnotherWonderfulMachine  26  qfwdsf
File1
YetAnother WonderfulMachine 3   sdg         
TrashWeWon'tBuy             4   jhfgjh          
MoreTrash                     5   qsfqf         
MiscelleneousStuff           6   qfsdf          
MoreMiscelleneousStuff       7   qsfwsf         
File2                   
MoreDifferentStuff          27  qsfsdf          
StrangeStuffBought          28  qsfsdf  

我在这里和那里尝试了一些 awks 脚本,但它们要么仅基于两个字段,我不知道如何修改输出,要么它们仅基于两个字段删除重复项,等等(我我是新手,awk 语法很难)。 非常感谢您的帮助。

【问题讨论】:

  • 文件 2 中的任何键都会在文件 1 中重复吗? IE。可能是多对多关系还是严格的一对多关系?在任何情况下,awk 可能不是用于此目的的最佳工具,它的数组运算符有点弱,您或多或少必须通过将两个文件读入大的关联数组然后对它们进行抨击来实现这一点。
  • 是的,两个文件中可能有重复的键。
  • 顺便说一句,如果您指定使用制表符来分隔输入和输出中的字段,则可以避免此问题引起的一定程度的混淆。

标签: awk merge


【解决方案1】:

您可以使用这三个命令非常接近:

join <(sort file1) <(sort file2)
join -v 1 <(sort file1) <(sort file2)
join -v 2 <(sort file1) <(sort file2)

这假定一个 shell,例如 Bash,支持进程替换 (&lt;())。如果您使用不支持的 shell,则需要对文件进行预排序。

在 AWK 中执行此操作:

#!/usr/bin/awk -f
BEGIN { FS="\t"; flag=1; file1=ARGV[1]; file2=ARGV[2] }
FNR == NR { lines1[$1] = $0; count1[$1]++; next }  # process the first file
{   # process the second file and do output
    lines2[$1] = $0;
    count2[$1]++;
    if ($1 != prev) { flag = 1 };
    if (count1[$1]) {
        if (flag) printf "%s ", lines1[$1];
        else printf "\t\t\t\t\t"
        flag = 0;
        printf "\t%s\n", $0
    }
    prev = $1
}
END { # output lines that are unique to one file or the other
    print "File 1: " file1
    for (i in lines1) if (! (i in lines2)) print lines1[i]
    print "File 2: " file2
    for (i in lines2) if (! (i in lines1)) print lines2[i]
}

运行它:

$ ./script.awk file1 file2

这些行的输出顺序与它们在输入文件中出现的顺序不同。第二个输入文件 (file2) 需要排序,因为脚本假定相似的行是相邻的。您可能需要调整脚本中的制表符或其他间距。我在这方面做得不多。

【讨论】:

  • join 方法将有效,即使两个输入文件中的键重复。在这种情况下,AWK 脚本只会从 file1 输出一组相似行中的最后一行(file2 中的重复项按预期处理)。
  • 谢谢!该脚本正在做正确的事情,但第一部分(将具有相同键的行放在同一行)以我无法解释的方式交错。它可能与字段分隔符(此处为制表符)有关。我有一些不幸的经历,试图在我发现做其他工作的脚本上修改字段分隔符。我发现 awk 的语法有点难以正确。(对不起,我不得不下机一个小时)。
  • @Trying:尝试添加else printf "\t\t\t\t\t",如我编辑的答案所示。
  • 我的第一个字段实际上是几个字段的串联,因此它读作“评论”类型的字段。合并是在字符串的第一个单词上完成的。
  • 我不确定我是否很清楚。字段如下所示:“这是一个句子作为第一个字段”选项卡“22”选项卡“其他字段”选项卡“又一个字段”
【解决方案2】:

一种方法(尽管使用硬编码的文件名):

BEGIN {
    FS="\t"; 
    readfile(ARGV[1], s1); 
    readfile(ARGV[2], s2); 
    ARGV[1] = ARGV[2] = "/dev/null"
}
END{
    for (k in s1) {
    if ( s2[k] ) printpair(k,s1,s2);
    }
    print "file1:"
    for (k in s1) {
    if ( !s2[k] ) print s1[k];
    }
    print "file2:"
    for (k in s2) {
    if ( !s1[k] ) print s2[k];
    }
}
function readfile(fname, sary) {
    while ( getline <fname ) {
    key = $1;
    if (sary[key]) {
        sary[key] = sary[key] "\n" $0; 
    } else {
        sary[key] = $0;
    };
    }
    close(fname);
}
function printpair(key, s1, s2) {
    n1 = split(s1[key],l1,"\n");
    n2 = split(s2[key],l2,"\n");
    for (i=1; i<=max(n1,n2); i++){
    if (i==1) {
        b = l1[1]; 
        gsub("."," ",b);
    }
    if (i<=n1) { f1 = l1[i] } else { f1 = b };
    if (i<=n2) { f2 = l2[i] } else { f2 = b };
    printf("%s\t%s\n",f1,f2);
    }
}
function max(x,y){ z = x; if (y>x) z = y; return z; }

不是特别优雅,但它可以处理多对多的情况。

【讨论】:

  • 您可以通过此更改将文件名作为命令行参数:BEGIN { readfile(ARGV[1], s1); readfile(ARGV[2], s2); ARGV[1] = ARGV[2] = "/dev/null" }
  • 它做对了,但输出没有考虑制表符分隔符,我不知道该放在哪里。
  • @Trying:标签?什么标签?!? ::看看丹尼斯回答中的 cmets:: 哦! 那个标签。好吧,你可以做丹尼斯做的同样的事情,替换 BEGIN 块中的字段分隔符; 在调用 readfile 之前这样做。同样,如果您想使用丹尼斯的巧妙小 ARGV 咒语。
  • 对不起,我修改了 BEGIN 块: "BEGIN { FS="\t"; readfile(ARGV[1], s1); readfile(ARGV[2], s2); ARGV [1] = ARGV[2] = "/dev/null" }" 我得到这个作为输出:“我的字段作为评论”TAB“22 第三个字段作为评论”TAB“36”,而不是“我的字段作为评论” TAB “22” “我的字段作为评论” TAB “36”。这意味着我无法抓取和排序我需要的那些数字。傻我,他?
  • @Trying:嗯...可能我们需要用制表符替换我用来衬托匹配线两侧的空格。会编辑。更改是printpair 中的printf 行。
猜你喜欢
  • 2014-03-31
  • 2023-04-03
  • 2015-01-25
  • 2019-01-14
  • 2016-01-12
  • 1970-01-01
  • 1970-01-01
  • 2019-06-14
  • 2019-03-01
相关资源
最近更新 更多