【问题标题】:Mathematically Subtracting Values in One File from Another File从另一个文件中数学减去一个文件中的值
【发布时间】:2019-07-26 05:11:56
【问题描述】:

我有两个文件,每个文件大小相同 (100x12),包含数值,正负均以逗号分隔。

文件 1 的示例输出:

-14.99,-15.6,8.0 ->
-9.0,34.87,98.98 ->
(and so on)

文件 2 的示例输出:

-15.99,-18.6,8.00 ->
-3.0,34.34,-98.88 ->
(and so on)

我试过了:

awk '{getline t<"file1"; print $0-t}' file2

但是,这只会减去第一列。如何扩展它以从 file2/column2 中减去 file1/column1?

我愿意使用 pandas 来执行此操作。提前谢谢!

【问题讨论】:

  • 对于这么小的问题,您不会从 pandas 中获得任何好处 - 您最好使用下面的解决方案 - 更快和更少的资源。 (而且我是熊猫的忠实粉丝!)
  • Re "从 file2/column2 中减去 file1/column1":这看起来很清楚,但它可能是一个错字。请确认给定样本中的第一个减法是-18.6 - -14.99 或-18.6 + 14.99(答案:3.61)。

标签: python bash pandas shell awk


【解决方案1】:

在我的脑海中 - 所以检查语法!

import numpy as np

with open("file1.txt, "r") as f1:
    with open("file2.txt, "r") as f2:
       array1 = np.asarray(f1.read().split(','))
       array2 = np.asarray(f2.read().split(','))
       result = array1 - array2
       print([x for x in result])

【讨论】:

    【解决方案2】:

    在 awk 中:

    $ awk '
    NR==FNR {                # hash file1 values to a
        for(i=1;i<=NF;i++)
            a[FNR][i]=$i
        next
    }{                       # process file2, subtract values from file1 respectives
        for(i=1;i<=NF;i++)
            $i=$i-a[FNR][i]
    }1' file1 file1
    

    输出:

    -1,-3,8.00 ->
    6,-0.53,-98.88 ->
    

    【讨论】:

      【解决方案3】:

      您可以尝试使用 unix 实用程序.. 使用 awk 粘贴

      paste file1.txt file2.txt | 
           awk -F"[,\t]"  -v OFS="," ' { for(i=1;i<4;i++) { $i=$(i+3)-$i } print $1,$2,$3 } '
      

      使用给定的输入

      $ cat halletwx1.txt
      -14.99,-15.6,8.0
      -9.0,34.87,98.98
      
      $ cat halletwx2.txt
      -15.99,-18.6,8.00
      -3.0,34.34,-98.88
      
      $ paste halletwx1.txt halletwx2.txt | awk -F"[,\t]"  -v OFS="," ' { for(i=1;i<4;i++) { $i=$(i+3)-$i } print $1,$2,$3 } '
      -1,-3,0
      6,-0.53,-197.86
      
      $
      

      【讨论】:

        【解决方案4】:

        首先是数据

        file1 = """
        -14.99,-15.6,8.0 ->
        -9.0,34.87,98.98 ->
        """
        file2 = """
        -15.99,-18.6,8.00 ->
        -3.0,34.34,-98.88 ->
        """
        
        from io import StringIO # faking file on disk
        

        熊猫回答。

        import pandas as pd
        converter = {2: lambda s: float(s.split(' ')[0])}
        df1 = pd.read_csv(StringIO(file1), header=None, converters=converter)
        df2 = pd.read_csv(StringIO(file2), header=None, converters=converter)
        (df1-df2).to_csv('pddiff12.csv', header=False, index=False)    
        

        或者用纯python滚动它。

        # cmt 1 -> indent under with-statement
        
        def read_csv(file_name):
            #with open('file_name', 'rt') as f1: # uncomment when reading from disk
            f1 = StringIO(file_name) # comment out when reading from disk
            rows = [r for r in f1.readlines() if r.strip()] # cmt 1
            crunch = lambda row: [float(r) for r in row.split(',')]
            rows = [crunch(r.split(' ')[0]) for r in rows]
            return rows
        
        data1 = read_csv(file1)
        data2 = read_csv(file2)
        
        diff = []
        for row1, row2 in zip(data1, data2):
            diff.append([i-j for i, j in zip(row1, row2)])
        
        with open('diff12.csv', 'wt') as d12:
            for row in diff:
                d12.write(', '.join((str(v) for v in row)) + '\n')
        

        Pandas 无疑是最容易阅读和工作的,尽管如果人们倾向于避免这种情况,这是一个明显的依赖关系。 在这种情况下,我想我不会。

        【讨论】:

        • 我选择使用 Pandas 解决方案;转换器的优雅解决方案 - 谢谢!
        • 老实说,从编写纯 python 开始认为 pandas 是不必要的臃肿,但在这种情况下,pandas 肯定也为我获得了 +1。
        猜你喜欢
        • 2018-01-25
        • 1970-01-01
        • 2021-12-22
        • 1970-01-01
        • 2016-09-28
        • 1970-01-01
        • 2021-04-11
        • 2013-02-21
        • 1970-01-01
        相关资源
        最近更新 更多