【问题标题】:Bash script to add double quotes in .CSV comma delimited file在 .CSV 逗号分隔文件中添加双引号的 Bash 脚本
【发布时间】:2020-07-27 21:55:23
【问题描述】:

我需要在 csv 文件中添加双引号。我的样本数据是这样的..

378478,COMPLETED,Tracfone,,,"2020/03/29 09:39:22",,2787,,356074101197544,89148000005748235454,75176540
378328,COMPLETED,"Total Wireless","Unlimited Talk, Text, & Data (First 25GB High Speed, then unlimited 2GB)",50,"2020/03/29 06:10:01",200890899011202395,0899,0279395,356058102052972,89148000005117597971,67756296

我已尝试使用awk 和sed 在线提供一些代码,结果如下,错误 - **数字中的第一个数字正在被修剪,如前。在“378478”中只显示“78478”。

此外,它还向已经存在的双引号添加双引号!** 似乎没有什么是完美的。请指导我!

"78478","COMPLETED","Tracfone","","",""2020/03/29 09:39:22"","","2787","","356074101197544","89148000005748235454","75176540"
"78328","COMPLETED",""Total Wireless"",""Unlimited Talk"," Text"," & Data (First 25GB High Speed"," then unlimited 2GB)"","50",""2020/03/29 06:10:01"","200890899011202395","0899","0279395","356058102052972","89148000005117597971","67756296"
"78329","COMPLETED",""Cricket Wireless"",""Unlimited Talk"," Text"," & 4G LTE Data w/ 15GB Hotspot"","60",""2020/03/29""

这是我正在使用的代码:

awk -F"'?,'?" -v OFS='","' '{$1=$1; gsub(/^.|$/,"\"")} 1' file # or
sed -E 's/([^,]*) , (.*)/"\1" , "\2"/' file

我的总代码如下。我的意图是首先将所有 .xlsx 转换为 .csv,然后将双引号添加到同一个 csv 并将其保存在同一个文件中。我知道 $file.csv 部分是错误的,因此我需要一些帮助

find "$Src_Dir" -type f -iname "*.xlsx" -print>path/temp

cat path/temp | while IFS="" read -r -d $'\0' file; 
do
    echo $file
    ssconvert "${file}" --export-type=Gnumeric_stf:stf_csv
    awk -F"'?,'?" -v OFS='","' '{$1=$1; gsub(/^.|$/,"\"")} 1' $file > $file.csv
done

【问题讨论】:

    标签: bash csv awk quotes comma


    【解决方案1】:

    如果您想处理除最简单 CSV 文件之外的任何内容,您可能应该将sed 和awk 移出。有更好的工具可用。

    例如,如果您在自己喜欢的发行版上使用sudo apt install csvtool(或同等版本),则可以使用其每行调用功能来处理输入文件中的每一行。有关示例,请参见以下脚本:

    #!/bin/bash
    
    function quotify {
      # Start empty line, process every field.
    
      line=""
      while [[ $# -ne 0 ]] ; do
          #    Append comma for all but first field, then quoted field.
    
          [[ -n "${line}" ]] && line="${line},"
          line="${line}\"$1\""
    
          shift
      done
    
      # Output the fully quoted line.
    
      echo "${line}"
    }
    
    # Needed to call functions. Also, ensure link: /bin/sh -> /bin/bash.
    export -f quotify
    
    # Pretty-print input and output.
    
    echo "Input file:"
    sed 's/^/   /' inputFile.csv
    
    echo "Output file:"
    csvtool call quotify inputFile.csv | sed 's/^/   /'
    

    注意 quotify 函数,该函数为 CSV 文件中的每个 line 调用,参数设置为该行中的每个 field(无引号,是否原始字段是否有引号)。

    它基本上构造了一行中所有字段的字符串,并在它们周围加上引号,然后将其写入标准输出,如下面该脚本的输出所示:

    Input file:
       378478,COMPLETED,Tracfone,,,"2020/03/29 09:39:22",,2787,,356074101197544,89148000005748235454,75176540
       378328,COMPLETED,"Total Wireless","Unlimited Talk, Text, & Data (First 25GB High Speed, then unlimited 2GB)",50,"2020/03/29"
    Output file:
       "378478","COMPLETED","Tracfone","","","2020/03/29 09:39:22","","2787","","356074101197544","89148000005748235454","75176540"
       "378328","COMPLETED","Total Wireless","Unlimited Talk, Text, & Data (First 25GB High Speed, then unlimited 2GB)","50","2020/03/29"
    

    即使使用单独的工具可能是最简单的方法,但如果您绝对无法安装其他软件包,那么您将不得不在已有的软件包中编写一些代码。以下bash 脚​​本是一个很好的起点,因为它不使用其他工具来实现其目标。

    目前,它与一组非常具体的规则相关联,如下所示:

    • 空白很重要。逗号之间的任何内容都被视为该字段的一部分。这在检测带引号的字段时尤其重要,它必须将引号作为第一个字符,没有 abc, "d,e,f",ghi 的东西,因为 "d,e,f" 将无法正确处理。
    • 引用字段允许包含逗号,其中的""序列转换为"。
    • 提供格式错误的 CSV 文件可能不是一个好主意 :-)

    但是,考虑到这一点,我们开始吧。我将提供每个部分的简短文字说明,但希望代码中的 cmets 足以弄清楚发生了什么。

    首先,一个用于查找某个字符串在另一个字符串中的位置的函数,用于计算字段边界:

    function findPos {
        haystack="$1"
        needle="$2"
    
        # Remove everything past the needle.
    
        prefix="${haystack%%${needle}*}"
    
        # If nothing was removed, it wasn't found, so supply massive number.
        # Otherwise, it was found at the length of the string with removed stuff.
    
        position=999999
        [[ ${#prefix} -ne ${#haystack} ]] && position=${#prefix}
        echo ${position}
    }
    

    然后我们可以在计算下一个字段长度的函数中使用它。这基本上只是为未引用的字段查找下一个逗号,并通过从段中构建字段来对引用的字段进行特殊处理(它必须处理引号和逗号中的引号):

    function getNextFieldLen {
        line="$1"
    
        # Empty line means all work done.
    
        [[ -z "${line}" ]] && echo -1 && return
    
        # Handle unquoted first, this is easy.
    
        [[ "${line:0:1}" != '"' ]] && { echo $(findPos "${line}" ","); return; }
    
        # Now handle quoted. Loop over all segments where a segment is defined as
        # the text up to the next <"">, assuming it's before the next <",>.
    
        field=""
        nextQuoteComma=$(findPos "${line}" '",')
        nextDoubleQuote=$(findPos "${line}" '""')
        while [[ ${nextDoubleQuote} -lt ${nextQuoteComma} ]]; do
            # Append segment to the field and go back for next segment.
    
            field="${field}${line:0:${nextDoubleQuote}}\"\""
            line="${line:${nextDoubleQuote}}"
            line="${line:2}"
    
            nextQuoteComma=$(findPos "${line}" '",')
            nextDoubleQuote=$(findPos "${line}" '""')
        done
    
        # Add final segment (up to the comma) and output entire field.
    
        field="${field}${line:0:${nextQuoteComma}}\""
        echo "${#field}"
    }
    

    最后,有一个顶级函数将引用通过标准输入输入的任何内容:

    function quotifyStdIn {
        # Process file line by line.
    
        while read -r line; do
            # Start with empty output line and non-comma separator.
    
            outLine="" ; sep=""
    
            # Place terminator to make processing easier, start field loop.
    
            line="${line},"
            fieldLen=$(getNextFieldLen "${line}")
            while [[ ${fieldLen} -ge 0 ]]; do
                # Get field and quotify if needed, adjust line (remove field and comma).
    
                field="${line:0:${fieldLen}}"
                [[ "${field:0:1}" = '"' ]] || field="\"${field}\""
    
                line="${line:$((fieldLen+1))}"
                #line="${line:${fieldLen}}"
                #line="${line:1}"
    
                # Append to output line and prepare for next field.
    
                outLine="${outLine}${sep}${field}"; sep=","
    
                fieldLen=$(getNextFieldLen "${line}")
            done
    
            # Output built line.
    
            echo "${outLine}"
        done
    }
    

    而且,如果您想直接从文件中读取数据(尽管提供空文件名或 "-" 将使用标准输入,因此您可能只使用基于文件的函数来处理所有内容):

    function quotifyFile {
        file="$1"
    
        # Empty file or "-" means standard input, otherwise take input from real file.
    
        [[ ${#file} -eq 0 ]] && { quotifyStdIn; return; }
        [[ "${file}" = "-" ]] && { quotifyStdIn; return; }
    
        quotifyStdIn < "${file}"
    }
    

    最后,因为每个程序不是“Hello, world”的程序都值得某种形式的测试工具,这就是您可以用来测试各种功能的工具:

    (
        echo 'paxdiablo,was here'
        echo 'and,"then, strangely,",he,was,not'
        echo '50,"My name is ""Pax"", and yours is ""Bob""",42'
        echo '17,"""Love"" is grand",19'
    ) > harness.csv
    
    echo "Before:"
    sed "s/^/   /" harness.csv
    echo "After:"
    quotifyFile harness.csv | sed "s/^/   /"
    
    rm -rf harness.csv
    

    而且,由于除非您运行测试,否则测试工具几乎没有用处,以下是第一次运行的结果:

    Before:
       paxdiablo,was here
       and,"then, strangely,",he,was,not
       50,"My name is ""Pax"", and yours is ""Bob""",42
       17,"""Love"" is grand",19
    After:
       "paxdiablo","was here"
       "and","then, strangely,","he","was","not"
       "50","My name is ""Pax"", and yours is ""Bob""","42"
       "17","""Love"" is grand","19"
    

    希望这足以让您在无法安装软件包的情况下继续前进。当然,如果您无法在 bash 本身中安装其中一个软件包,那么您遇到的问题我无法帮助您解决:-)

    【讨论】:

    • 好的,这很好。我以前没有使用过 ocaml-csv。当然让事情变得更容易。
    • @Alekhyavarma,你可以这样做( cd /path ; mv filename.csv filename.csv.bak &amp;&amp; quotifyFile filename.csv.bak &gt; filename.csv ) - 这应该会给你一个修改后的文件,并使用原始名称并为你留下一个原始文件的备份,以防万一。像cat myfile &gt;myfile 这样的命令将不起作用,因为重定向(截断myfile)发生在shell 之前 cat 命令试图打开它。因此,那时它已经是空的了。
    • @paxdiablo - 宾果游戏!!好男人!!你是真正的救世主!!它工作得非常完美!
    【解决方案2】:

    您的起始 CSV 不是一个好的 CSV:2 行的列数不同

    +--------+-----------+----------------+--------------------------------------------------------------------------+----+---------------------+---+------+---+-----------------+----------------------+----------+
    | 1      | 2         | 3              | 4                                                                        | 5  | 6                   | 7 | 8    | 9 | 10              | 11                   | 12       |
    +--------+-----------+----------------+--------------------------------------------------------------------------+----+---------------------+---+------+---+-----------------+----------------------+----------+
    | 378478 | COMPLETED | Tracfone       | -                                                                        | -  | 2020/03/29 09:39:22 | - | 2787 | - | 356074101197544 | 89148000005748235454 | 75176540 |
    | 378328 | COMPLETED | Total Wireless | Unlimited Talk, Text, & Data (First 25GB High Speed, then unlimited 2GB) | 50 | 2020/03/29          | - | -    | - | -               | -                    | -        |
    +--------+-----------+----------------+--------------------------------------------------------------------------+----+---------------------+---+------+---+-----------------+----------------------+----------+
    

    使用 Miller (https://github.com/johnkerl/miller) 你可以运行

    mlr --csv --quote-all -N unsparsify input >output
    

    拥有

    "378478","COMPLETED","Tracfone","","","2020/03/29 09:39:22","","2787","","356074101197544","89148000005748235454","75176540"
    "378328","COMPLETED","Total Wireless","Unlimited Talk, Text, & Data (First 25GB High Speed, then unlimited 2GB)","50","2020/03/29","","","","","",""
    

    您可以使用它下载可执行文件https://github.com/johnkerl/miller/releases/tag/v5.7.0

    【讨论】:

      猜你喜欢
      • 2016-07-03
      • 1970-01-01
      • 1970-01-01
      • 2014-10-03
      • 1970-01-01
      • 1970-01-01
      • 2017-09-28
      • 2023-03-21
      • 1970-01-01
      相关资源
      最近更新 更多