【问题标题】:Split comma separated list with line number using awk and bash使用 awk 和 bash 用行号拆分逗号分隔列表
【发布时间】:2020-02-04 14:42:47
【问题描述】:

我有一个(非常大的)csv 文件,格式如下:

id;surname;firstname;aliases
1;Simpson;Homer;Homer Jay Simpson,Homer J. Simpson
2;Simpson;Bart;Bartholomew JoJo Simpson,Bartholomew Simpson
3;Krusty the Clown;;Herschel Shmoikel Pinchas Yerucham Krustofsky
4;Simpson;Lisa;

现在我想将其转换为以下格式:

id;name
1;Homer Simpson
1_1;Homer Jay Simpson
1_2;Homer J. Simpson
2;Bart Simpson
2_1;Bartholomew JoJo Simpson
2_2;Bartholomew Simpson
3;Krusty the Clown
3_1;Herschel Shmoikel Pinchas Yerucham Krustofsky
4;Lisa Simpson

出于性能原因,我想使用 awk 或其他 UNIX 命令行工具来实现。

使用awk -F ';' '{print $1, $3, $2}' 我可以分隔分号分隔的行。但是如何在awk 中使用awk 再次拆分逗号分隔的条目?

【问题讨论】:

  • 你已经包含了一个有姓但没有名字的案例(Krusty the Clown),如果你可以有相反的,那么你也应该在你的例子中包含它。

标签: shell awk


【解决方案1】:
$ cat tst.awk
BEGIN { FS=OFS=";" }
NR==1 {
    print $1, "name"
    next
}
{
    name = $3 " " $2
    gsub(/^ +| +$/,"",name)
    print $1, name
    n = split($NF,aliases,/,/)
    for (i=1; i<=n; i++) {
        print $1 "_" i, aliases[i]
    }
}

$ awk -f tst.awk file
id;name
1;Homer Simpson
1_1;Homer Jay Simpson
1_2;Homer J. Simpson
2;Bart Simpson
2_1;Bartholomew JoJo Simpson
2_2;Bartholomew Simpson
3;Krusty the Clown
3_1;Herschel Shmoikel Pinchas Yerucham Krustofsky
4;Lisa Simpson

【讨论】:

    【解决方案2】:

    您能否尝试以下操作(使用所示示例编写和测试)。

    awk '
    BEGIN{
      FS="[;,]"
      OFS=";"
      print "id;name"
    }
    FNR>1{
      j=$2~/ /?2:3
      for(i=j;i<=NF;i++){
        if($i==""){
          continue
        }
        if(i==j){
          print $1,$3" "$2
        }
        else{
          print $1"_"++c,$i
        }
      }
      c=""
    }' Input_file
    

    输出如下。

    id;name
    1;Homer Simpson
    1_1;Homer Jay Simpson
    1_2;Homer J. Simpson
    2;Bart Simpson
    2_1;Bartholomew JoJo Simpson
    2_2;Bartholomew Simpson
    3; Krusty the Clown
    3_1;Herschel Shmoikel Pinchas Yerucham Krustofsky
    4;Lisa Simpson
    

    说明:在此添加上述代码的详细说明。

    awk '                        ##Starting awk program from here.
    BEGIN{                       ##Starting BEGIN section from here.
      FS="[;,]"                  ##Setting field as either semi-colon OR comma for all lines.
      OFS=";"                    ##Setting output field separator semi-colon.
      print "id;name"            ##Printing id;name string before reading Input_file.
    }                            ##Closing BLOCK for BEGIN block of this awk program here.
    FNR>1{                       ##Checking condition if FNR>1 then do following.
      j=$2~/ /?2:3
      for(i=j;i<=NF;i++){        ##Running a for loop from i=j to till number of fields of line.
        if($i==""){              ##Checking condition if current field value is NULL then do following.
          continue               ##Using continue to take cursor to for loop again here.
        }
        if(i==j){                ##Checking condition if i==3 then do following.
          print $1,$3" "$2       ##Printing first, 3rd,space and 2nd field of line here.
        }
        else{                    ##If above if condition is false then come to this else here.
          print $1"_"++c,$i      ##Printing first field underscore variable c value, value of current field here.
        }
      }
      c=""                       ##Nullifying variable c here.
    }
    '  Input_file                ##Mentioning Input_file name here.
    

    【讨论】:

    • 小丑克鲁斯蒂失踪了;这是一个错误吗?
    • @tripleee,谢谢先生,我已经修好了。
    【解决方案3】:

    这将按照您在 Python 3 中的意图进行。请注意,我键入它的速度很快,因此可以进行许多改进。我相信它可能比 awk 快,但我可能错了。您可以在 Linux 和 Mac 中使用 time 命令测试是否如此。

    #!/usr/local/bin/python3
    
    import csv
    csvr = csv.reader(open('simpsons.csv'), delimiter = ";")
    
    index=0
    for row in csvr:
        if index == 0:
            index = index +1
            continue
        print("{};{} {}".format(index,row[2],row[1]))
        sindex=0
        for sitem in row[3].split(','):
            if sitem != "" :
                sindex = sindex + 1
                print("{};{}".format(row[0] + "_" + str(sindex),sitem))
        index = index +1
    

    希望对你有帮助!

    编辑:

    我生成了一个 500k 行的虚拟列表,并测试了用户在这里给出的一些答案,这似乎不是 Python 3 和 awk 之间的任何重要区别。 (至少在我在 Python 3 中的糟糕实现中)。

     $ time awk -f tst.awk fivehundredthousand.txt &> /dev/null
    
    real    0m2.141s
    user    0m2.118s
    sys     0m0.020s
    
     $ time ./handle_csv.py >/dev/null
    
    real    0m1.750s
    user    0m1.722s
    sys     0m0.021s
    
    $ time awk -f ravinder.awk fivehundredthousand.txt &> /dev/null
    
    real    0m1.736s
    user    0m1.718s
    sys     0m0.017s
    

    【讨论】:

    • 对于您的计时结果 - 每个脚本的第三次运行计时是否可以消除缓存的影响?
    • @EdMorton 实际上我每个只运行一次。在这种情况下是否有缓存?
    • 是的,由于缓存,后续运行可能比初始运行更快,因此您必须始终运行任何命令 3 次,然后再将第 3 次执行时间与任何其他命令的时间(您'd 也会运行 3 次)。
    • 太棒了。我不知道有进程级缓存。跟CPU缓存有关系吗?你有与此相关的链接吗?
    • Idk 它与什么有关,我只是在谷歌上搜索参考,但找不到,抱歉。我可能会再尝试几分钟的谷歌搜索,但查询起来很困难,因为它会产生很多不相关的点击。在unix.stackexchange.com/q/8398/133219FWIW 上提到了这种效果。
    【解决方案4】:

    Awk 有一个split 函数,可以让您将字符串拆分为数组。

    awk -F ';' 'BEGIN { OFS=FS }
      { print $1, $3 " " $2
        n = split($4, alias, /,/)
        for(i=1; i<=n; i++)
          print $1 "_" i, alias[i] }' file.csv
    

    split 的返回值告诉您结果数组中有多少成员。

    【讨论】:

    • 如果您想摆脱 Krusty 之前的空间,请尝试 ($3 ? $3 " " : "")。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2013-07-03
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-03-07
    • 1970-01-01
    相关资源
    最近更新 更多