【问题标题】:Creating multiple csv files from data within a csv file从 csv 文件中的数据创建多个 csv 文件
【发布时间】:2011-02-06 23:18:28
【问题描述】:

系统 OSX 或 Linux

我正在尝试自动化我的工作流程,每周我都会收到一个 excel 文件,然后将其转换为 csv。

一个例子是:

,,L1,,,L2,,,L3,,,L4,,,L5,,,L6,,,L7,,,L8,,,L9,,,L10,,,L11,
Title,r/t,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst,neede d,actual,Inst,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst,needed,actual,Inst
EXAMPLEfoo,60,6,6,6,0,0,0,0,0,0,6,6,6,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0
EXAMPLEbar,30,6,6,12,6,7,14,6,6,12,6,6,12,6,8,16,6,7,14,6,7.5,15,6,6,12,6,8,16,6,0,0,6,7,14
EXAMPLE1,60,3,3,3,3,5,5,3,4,4,3,3,3,3,6,6,3,4,4,3,3,3,3,4,4,3,8,8,3,0,0,3,4,4
EXAMPLE2,120,6,6,3,0,0,0,6,8,4,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0
EXAMPLE3,60,6,6,6,6,8,8,6,6,6,6,6,6,0,0,0,0,0,0,6,8,8,6,6,6,0,0,0,0,0,0,0,10,10
EXAMPLE4,30,6,6,12,6,7,14,6,6,12,6,6,12,3,5.5,11,6,7.5,15,6,6,12,6,0,0,6,9,18,6,0,0,6,6.5,13

这样您就可以了解它在 excel 中的外观:

我需要做的是为第 1 行中的每个实例创建多个 csv 文件,因此 L1、L2、L3、L4...

在每个 csv 文件中,它需要包含标题,r/t,需要

因此对于 L1,输出示例如下所示:

EXAMPLEfoo,60,6
EXAMPLEbar,30,6
EXAMPLE1,60,3
EXAMPLE2,120,6
EXAMPLE3,60,6
EXAMPLE4,30,6

对于 L2:

EXAMPLEfoo,60,0
EXAMPLEbar,30,6
EXAMPLE1,60,3
EXAMPLE2,120,0
EXAMPLE3,60,6
EXAMPLE4,30,6

等等。

我尝试过使用 sed 和 awk 并点击 google,但我没有找到任何真正解决问题的方法。

我想 perl 会特别适合这个或 python,所以我很乐意接受用户的建议。

那么,有什么建议吗?

提前致谢。

【问题讨论】:

  • 您收到了哪些类型的 Excel 文件——2003 年、2007 年,其他?
  • 2007 在 mac 上,我已经搜索了尝试使用 automator 执行此操作的方法,它是 excel 挂钩,但没有骰子。理想情况下,我希望能够从 bash 脚本运行它,这解释了我对 sed 和 awk 的实验。
  • 看看我的 FOSS 项目code.google.com/p/csvfix,它是一个可以做这类事情的工具。
  • 鉴于各种L的覆盖多列,你如何确定选择哪一个?
  • 嗨,第一行的条目数是常数吗? L1 是唯一一个错过 Inst 列的吗?

标签: python perl bash sed awk


【解决方案1】:

Perl“单行”

perl -MText::CSV_XS -e'$c=Text::CSV_XS->new({binary=>1,eol=>"\n"});%a=map{$i++;/^L\d+$/?($_=>$i):()}@{$c->getline(*ARGV)};open$b{$_},">$_"for keys%a;while($f=$c->getline(*ARGV)){$c->print($b{$_},[@$f[0,1,$a{$_}]])for keys%a}'

对于阅读有问题的人:

$ echo '$c=Te...' | perltidy
$c = Text::CSV_XS->new( { binary => 1, eol => "\n" } );
%a = map { $i++; /^L\d+$/ ? ( $_ => $i ) : () } @{ $c->getline(*ARGV) };
open $b{$_}, ">$_" for keys %a;
while ( $f = $c->getline(*ARGV) ) {
    $c->print( $b{$_}, [ @$f[ 0, 1, $a{$_} ] ] )
      for keys %a;
}

【讨论】:

    【解决方案2】:

    仅使用 AWK:

    awk -F, -vOFS=, -vc=1 '
        NR == 1 {
            for (i=1; i<NF; i++) {
                if ($i != "") {
                    g[c]=i;
                    f[c++]=$i
                }
            }
        }
        NR>2 {
            for (i=1; i < c; i++) {
                print $1,$2, $g[i] > "output_"f[i]".csv"
            }
        }' data.csv
    

    作为单行:

    awk -F, -vOFS=, -vc=1 'NR == 1 {for (i=1; i<NF; i++) {if ($i != "") {g[c]=i; f[c++]=$i}}} NR>2 { for (i=1; i < c; i++) {print $1,$2, $g[i] > "file_"f[i]".csv" }}' data.csv
    

    示例输出:

    $ cat file_L1.csv
    EXAMPLEfoo,60,6
    EXAMPLEbar,30,6
    EXAMPLE1,60,3
    EXAMPLE2,120,6
    EXAMPLE3,60,6
    EXAMPLE4,30,6
    $ cat file_L2.csv
    EXAMPLEfoo,60,0
    EXAMPLEbar,30,6
    EXAMPLE1,60,3
    EXAMPLE2,120,0
    EXAMPLE3,60,6
    EXAMPLE4,30,6
    $ cat file_L11.csv
    EXAMPLEfoo,60,0
    EXAMPLEbar,30,6
    EXAMPLE1,60,3
    EXAMPLE2,120,0
    EXAMPLE3,60,0
    EXAMPLE4,30,6
    

    【讨论】:

    • 我没有得到 OP 所述的输出。
    • 糟糕,我忘记拿出测试打印件了。
    • NR == 1 循环构建包含L1 等的单元格的位置和内容的数组。NR&gt;2 循环在每个记录中水平移动,将正确的数据输出到正确的文件.在我的系统上似乎附加了&gt;,也许重定向应该更改为&gt;&gt;。您的脚本遍历整个文件 11 次。我的只做一次。
    • 感谢您的回复,我希望最终会与每个人联系,但是丹尼斯,您的看起来很有希望,所以我将从这里开始。您在哪个版本的 awk 上运行它? $ awk -F, -vOFS=, -vc=1 'NR == 1 {for (i=1; i2 { for (i=1; i "file_"f[i]".csv" } }' data.csv awk: invalid -v option 我将尝试从默认的 OSx 安装升级,即 awk 版本 20070501,我会发回我的结果。
    • @S1syphus:尝试将其更改为-v OFS=","(带有空格和引号)。我正在使用 GNU AWK (gawk) 3.1.6。
    【解决方案3】:
    use strict;
    use warnings;
    
    use Text::CSV;
    my $csv = Text::CSV->new;
    
    sub parse_line {
        $csv->parse(shift) or die $!;
        return $csv->fields;
    }
    
    my @metadata;
    my @files  = parse_line(scalar <>);
    my @header = parse_line(scalar <>); # Ignore.
    for my $i (0 .. $#files){
        next unless length $files[$i];
        open(my $h, '>', "$files[$i].csv") or die $!;
        push @metadata, {column => $i, handle => $h};
    }
    
    while (my $line = <>){
        my @fields = parse_line($line);
        for my $m (@metadata){
            $csv->print($m->{handle}, [ @fields[0, 1, $m->{column}] ]);
            print {$m->{handle}} "\n";
        }
    }
    

    【讨论】:

      【解决方案4】:

      试试这个

      #!/bin/bash
      awk 'BEGIN{ OFS=FS="," }
      NR==1{
       for(i=1;i<=NF;i++){
         if($i){ f[i]=$i }
       }
      }
      NR>2{ for(o in f){ print $1,$2, $o > "file_"f[o]".csv" } } ' file
      

      输出

      $ cat file_L1.csv
      EXAMPLEfoo,60,6
      EXAMPLEbar,30,6
      EXAMPLE1,60,3
      EXAMPLE2,120,6
      EXAMPLE3,60,6
      EXAMPLE4,30,6
      
      $ cat file_L2.csv
      EXAMPLEfoo,60,0
      EXAMPLEbar,30,6
      EXAMPLE1,60,3
      EXAMPLE2,120,0
      EXAMPLE3,60,6
      EXAMPLE4,30,6
      

      【讨论】:

      • 感谢您的回复,我想这可能只是我,因为没有一个像这样完成的示例,包括你的示例给了我一个错误,类似于 awk:source line 7 context 的语法错误是NR>2{ for(o in f){ print $1,$2, $o > >>> "file_"f
      • 我不知道为什么。您可以尝试使用 dos2unix(或 unix2dos)将脚本的换行符转换为适合 OSX 的换行符。
      【解决方案5】:

      查看 perl 模块 Text::CSV_XS - 逗号分隔值操作例程。我发现这个模块在处理 CSV 文件时非常有用。

      【讨论】:

        【解决方案6】:

        在 Python 中,略显老套且未经测试,但应该可以胜任:

        import csv
        r = csv.reader(open(r'file.csv'), dialect='excel')
        topline = r.next()
        headerline = r.next()
        
        lastcell = ''
        for i, cell in enumerate(topline): #Copy cells forwards in the top line, so L1 for example goes across all cells
            if cell == '':
                topline[i] = lastcell
            else:
                lastcell = cell
        
        for i in range(len(headerline)): #Copy the topline cells into the header line, so the headerline cells should be unique
            headerline[i] = '-'.join((topline[i], headerline[i]))
        
        rows = [dict(zip(headerline, line)) for line in r]
        
        # Rows should now consist of dicts of the form {'Title': 'EXAMPLEfoo', 'r/t': '60', 'L1-needed': '6' ...}
        
        for lval in frozenset(topline): #Use frozenset to ensure we only have unique values.
            if lval != '': #Make sure we don't look at the blank value
                w = csv.writer(open(r'%s.csv' % lval, 'w'), dialect='excel')
                for row in rows:
                    line = [row['Title'], row['r/t'], row['-'.join((lval, 'needed'))]]
                    w.writerow(line)
        

        【讨论】:

        • 您的“for cell in topline:”循环实际上并没有做任何事情;分配给“cell”不会改变“topline”的元素。您需要将其替换为 for i, cell in enumerate(topline): if cell == '': topline[i] = lastcell 等等。为什么你在最终循环中使用frozenset,而不是仅仅设置?
        • @Peter:好地方;我总是忘记 Python 列表枚举中的那个缝隙。我会及时更新。我正在使用frozenset,因为当时它似乎是个好主意...... AFAIK 在这种情况下没有任何区别。
        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2019-03-24
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2018-05-09
        • 2011-03-19
        • 2021-09-17
        相关资源
        最近更新 更多