【问题标题】:Counting letters in a file in shell script在shell脚本中计算文件中的字母
【发布时间】:2015-08-03 02:06:15
【问题描述】:

我需要一个 shell 脚本/powershell,用来计算文件中相似的字母。

输入:

this is the sample of this script.
This script counts similar letters.

输出:

t 9
h 4
i 8
s 10
e 4
a 2
...

【问题讨论】:

    标签: shell powershell scripting


    【解决方案1】:

    这个班轮应该做的:

    awk  'BEGIN{FS=""}{for(i=1;i<=NF;i++)if(tolower($i)~/[a-z]/)a[tolower($i)]++}
          END{for(x in a)print x, a[x]}' file
    

    您的示例的输出:

    u 1
    h 4
    i 8
    l 3
    m 2
    n 1
    a 2
    o 2
    c 3
    p 3
    r 4
    e 4
    f 1
    s 10
    t 9
    

    【讨论】:

    • @aphoria 我看到shell script/powershell 你将这一行保存在一个文件中,那么它就是 shell 脚本。还有带有shell标签的问题在该标签上移动鼠标,将显示解释。
    • 再说一次,我没有对你投反对票,所以你不必向我解释……我只是在猜测为什么其他人可能对你投了反对票。另一个不是 PowerShell 解决方案的答案也被否决了。我并不是说你应该被否决,这只是一个可能的原因。
    【解决方案2】:

    在 PowerShell 中,您可以使用 Group-Object cmdlet:

    function Count-Letter {
        param(
            [String]$Path,
            [Switch]$IncludeWhitespace,
            [Switch]$CaseSensitive
        )
    
        # Read the file, convert to char array, and pipe to group-object
        # Convert input string to lowercase if CaseSensitive is not specified
        $CharacterGroups = if($CaseSensitive){
            (Get-Content $Path -Raw).ToCharArray() | Group-Object -NoElement
        } else {
            (Get-Content $Path -Raw).ToLower().ToCharArray() | Group-Object -NoElement
        }
    
        # Remove any whitespace character group if IncludeWhitespace parameter is not bound
        if(-not $IncludeWhitespace){
            $CharacterGroups = $CharacterGroups |Where-Object { "$($_.Name)" -match "\S" }
        }
    
        # Return the groups, letters first and count second in a default format-table
        $CharacterGroups |Select-Object @{Name="Letter";Expression={$_.Name}},Count
    }
    

    这是我的机器上的输出与您的示例输入 + 换行符的样子

    【讨论】:

    • 谢谢,但它在 PS1 格式中的外观如何?我需要这样的输入:task.ps1 letters.txt
    • @MolnárBence 删除function Count-Letter {} 块,以便task.ps1 文件中的第一行是param( 开头 - 然后您可以像您描述的那样调用它。如果您不希望标题 + 分隔符位于输出的顶部,请通过管道将其发送到 Format-Table -HideTableHeaders
    • 我删除了块,但它不起作用。可以看到问题:people.inf.elte.hu/bencehun93/error.jpg
    • @MolnárBence 正如错误所述,它与代码本身无关,而与您机器上的执行策略设置有关。如果您发现自己无法使用 Set-ExecutionPolicy 更改它,请尝试像这样启动 PowerShell:powershell.exe -executionpolicy bypass 并从那里尝试
    【解决方案3】:

    powershell 一班:

    "this is the sample of this script".ToCharArray() | group -NoElement | sort Count -Descending | where Name -NE ' '
    

    【讨论】:

    • 我会在排序之前将过滤器向上移动(不需要对你要丢弃的东西进行排序)
    【解决方案4】:
    echo "this is the sample of this script"  | \
    sed -e 's/ //g' -e 's/\([A-z]\)/\1|/g'  |  tr '|' '\n'  |  \
    sort  |  grep -v "^$"  |  uniq -c  |  \
    awk '{printf "%s %s\n",$2,$1}'
    

    【讨论】:

      【解决方案5】:
      echo "this is the sample of this script. \
      This script counts similar letters." | \
          grep -o '.' | sort | uniq -c | sort -rg
      

      首先输出、排序、最常见的字母:

       10 s
       10  
        8 t
        8 i
        4 r
        4 h
        4 e
        3 p
        3 l
        3 c
        2 o
        2 m
        2 a
        2 .
        1 u
        1 T
        1 n
        1 f
      

      注意:不需要sed 或awk;一个简单的grep -o '.' 完成所有繁重的工作。要不计算空格和标点符号,请将'.' 替换为'[[:alpha:]]' |:

      echo "this is the sample of this script. \
      This script counts similar letters." | \
          grep -o '[[:alpha:]]' | sort | uniq -c | sort -rg
      

      要将大小写字母算作一个,请使用sort 和uniq 的--ignore-case 选项:

      echo "this is the sample of this script. \
      This script counts similar letters." | \
          grep -o '[[:alpha:]]' | sort -i | uniq -ic | sort -rg
      

      输出:

       10 s
        9 t
        8 i
        4 r
        4 h
        4 e
        3 p
        3 l
        3 c
        2 o
        2 m
        2 a
        1 u
        1 n
        1 f
      

      【讨论】:

        猜你喜欢
        • 2023-03-31
        • 2011-06-28
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2021-03-27
        • 2021-09-01
        • 2010-11-25
        相关资源
        最近更新 更多