【问题标题】:merge two files based on partial matching基于部分匹配合并两个文件
【发布时间】:2019-11-18 11:07:39
【问题描述】:

我有两个文件

文件A.txt

ID
479432_Sros_4274
330214_NIDE2792
517722_CJLT1_010100003977
257310_BB0482
...

FileB.txt(**只是为了帮助您识别匹配项)

members   category
6085.XP_002168109,**479432_Sros_4274**,4956.XP_002495993.1,457425.SSHG_03214,51511.ENSCSAVP000  P
7159.AAEL006372-PA,**257310_BB0482** J
**517722_CJLT1_010100003977**,701176.VIBRN418_17773,9785.ENSLAFP00000010769,28377.ENSACAP00000014901,4081.Solyc03g120250.2.1,3847.GLYMA18G02240.1 U
500485.XP_002561312.1,1042876.PPS_0730,222929.XP_003071446.1,**330214_NIDE2792**  S
...

预期输出

输出.txt

ID  category
479432_Sros_4274  P
330214_NIDE2792  S
517722_CJLT1_010100003977  U
257310_BB0482  J
...

我根据其他问题的答案在 awk 和 R 中尝试了一些代码,但我无法获得所需的输出。

【问题讨论】:

    标签: bash unix awk merge match


    【解决方案1】:

    这是一种方法:

    $ awk '
    NR==FNR {                  # process file1
        if(FNR==1)             # print header, no newline
            printf $1
        a[$1]                  # hash data
        next
    }
    {                          # process file2
        if(FNR==1)             # print the other half of the header
            print OFS $2
        for(i in a)            # loop all items in hash
            if($1 ~ i)         # check for partial match
                print i,$2     # if found, output
    }' file1 file2             # mind the order
    

    输出(按file2顺序,注意输出最后一行的部分匹配,留作警告):

    ID category
    479432_Sros_4274 P
    257310_BB0482 J
    517722_CJLT1_010100003977 U
    330214_NIDE2792 S
    ID S
    

    【讨论】:

    • 感谢您的回答,但输出不是预期的。输出看起来像:``` ID 类别 ID P ID S ID U ID J ... ``` 似乎 ID 在 file2 中出现了很多次(我用 grep 检查过)。然后我更改了 file2 中不存在的“name_of_identifier”的列标题 ID,然后输出文件只包含 1 行的标题:```name_of_identifier 类别```
    • 啊,我刚刚仔细检查了一遍。我不得不替换“。” 1 个文件中的“_”与另一个文件匹配。现在效果很好!谢谢!
    • 似乎 ID 确实在 file2 中出现了很多次 是的,这就是为什么在答案中将其作为警告而不是修复它的原因。简单的解决方法是在printf 之后添加next,即。 if(FNR==1){printf $1; next} 或像@RavinderSingh13 一样跳过BEGIN 块中的FNR==1 和print 标头。
    【解决方案2】:

    请您尝试关注一下。

    awk '
    BEGIN{
      print "ID  category"
    }
    FNR==NR{
      a[$0]
      next
    }
    {
      for(i in a){
        if(match($0,i)){
          print i,$NF
        }
      }
    }
    '  Input_filea   Input_fileb
    

    说明:为上述代码添加说明。

    awk '                               ##Starting awk program here.
    BEGIN{                              ##Starting BEGIN section from here.
      print "ID  category"              ##Printing string ID, category here.
    }                                   ##Closing BLOCK for BEGIN section.
    FNR==NR{                            ##Checking condition FNR==NR which will be TRUE when 1st Input_file is being read.
      a[$0]                             ##Creating an array named a whose index is $).
      next                              ##next will skip all further statements from here.
    }
    {
      for(i in a){                      ##Traversing through array a with for loop.
        if(match($0,i)){                ##Checking condition if match is having a proper regex matched then do following.
          print i,$NF                   ##Printing variable i and $NF of current line.
        }
      }
    }
    '  Input_filea   Input_fileb        ##Mentioning Input_file names here.
    

    【讨论】:

    • 感谢您的回答,但输出不是预期的。输出看起来像:``` ID 类别 ID P ID S ID U ID J ... ``` 似乎 ID 在 file2 中出现了很多次(我用 grep 检查过)。然后我更改了 file2 中不存在的“name_of_identifier”的列标题 ID,然后输出文件只包含 1 行的标题:```name_of_identifier 类别```
    • @palomo11,我的代码是使用提供的示例编写和测试的,如果您的数据与显示的数据不同,请使用更接近您的数据的数据编辑您的帖子,一旦编辑完成,请让我知道。
    • 啊,我刚刚仔细检查了一遍。我不得不替换“。” 1 个文件中的“_”与另一个文件匹配。现在效果很好!谢谢!
    猜你喜欢
    • 2020-10-15
    • 1970-01-01
    • 1970-01-01
    • 2015-08-20
    • 2021-01-29
    • 1970-01-01
    • 1970-01-01
    • 2013-03-28
    相关资源
    最近更新 更多