【问题标题】:How to extract data from a text file in Python?如何从 Python 中的文本文件中提取数据?
【发布时间】:2019-06-25 07:28:22
【问题描述】:

我有一个包含大量信息的文本文件。我试图为此寻求一些帮助。我发现了一些有点相似但不完全是我想做的事情。

我有一个文本文件(如下所示),我想从中提取前 3 列的数据到一个数组中。

我是 Python 的初学者。请帮助解决这个问题。

//Text file starts
---------------------------SOFTWARE NAME------------------------------------

I/O Filenames:  abc.txt           
Variables:______                       
------------------------------------------------------------------------

 Method name.

  Coordinates :      0         0

     S.No.      X(No.)    Y(No.)      Z(Beta)    A(Alpha)

     1    3.541            0
     2    7.821          180
     3    2.160            0
     4    4.143            0    3.69            0
     5    2.186            0    2.18            0
     6    3.490            0    2.45            0
//End of text file

【问题讨论】:

  • 作为初学者并不是不尝试任何事情的理由。您将通过编写一些代码然后寻求帮助来了解更多信息。提示:由于该文件不是真正的 csv 文件,我的建议是:1/逐行读取文件 2/剥离第 3 行/忽略任何不以数字开头的行(strip 之后)3/拆分该行并保留前 3 个字段。

标签: python python-3.x file csv


【解决方案1】:

为此,我将使用包csv

import csv #Import the package

with open('/path/to/file.csv') as csvDataFile: #open the csv file
    csvReader = csv.reader(csvDataFile,delimiter=';') #load the csv file with the delimiter of your choice, here it is a ; 
    for row in csvReader:
      #do something with the row

我建议你更好地格式化你的文件,一个好的是:

S.No.;X(No.);Y(No.);Z(Beta);A(Alpha)
1;3.541;0;;
2;7.821;180;;
3;2.160;0;;
4;4.143;0;3.69;0
5;2.186;0;2.18;0
6;3.490;0;2.45;0

这里是a link 了解更多信息

【讨论】:

    【解决方案2】:

    您可以尝试使用re 模块(regex101)提取数据:

    import re
    from itertools import zip_longest
    
    data = '''
    //Text file starts
    ---------------------------SOFTWARE NAME------------------------------------
    
    I/O Filenames:  abc.txt
    Variables:______
    ------------------------------------------------------------------------
    
     Method name.
    
      Coordinates :      0         0
    
         S.No.      X(No.)    Y(No.)      Z(Beta)    A(Alpha)
    
         1    3.541            0
         2    7.821          180
         3    2.160            0
         4    4.143            0    3.69            0
         5    2.186            0    2.18            0
         6    3.490            0    2.45            0
    //End of text file
    '''
    
    l = [g.split() for g in re.findall(r'^\s+\d+\s+[^\n]+$', data, flags=re.M)]
    for v in zip(*zip_longest(*l)):
        print(v)
    

    打印:

    ('1', '3.541', '0', None, None)
    ('2', '7.821', '180', None, None)
    ('3', '2.160', '0', None, None)
    ('4', '4.143', '0', '3.69', '0')
    ('5', '2.186', '0', '2.18', '0')
    ('6', '3.490', '0', '2.45', '0')
    

    【讨论】:

    【解决方案3】:

    使用re 模块提取文本。使用numpy 模块来构造你需要的“数组”。

    import re 
    import numpy as np
    
    text = """
    //Text file starts
    ---------------------------SOFTWARE NAME------------------------------------
    
    I/O Filenames:  abc.txt           
    Variables:______                       
    ------------------------------------------------------------------------
    
     Method name.
    
      Coordinates :      0         0
    
         S.No.      X(No.)    Y(No.)      Z(Beta)    A(Alpha)
    
         1    3.541            0
         2    7.821          180
         3    2.160            0
         4    4.143            0    3.69            0
         5    2.186            0    2.18            0
         6    3.490            0    2.45            0
    //End of text file
    """
    
    regex = r'(?<=^\s{5})\d\s*[\d\.]*\s*\d*'
    matches = [x.split() for x in re.findall(regex, text, flags=re.MULTILINE)]
    
    arr = np.array(matches)
    print(arr)
    

    它提供输出:

    [['1' '3.541' '0']
     ['2' '7.821' '180']
     ['3' '2.160' '0']
     ['4' '4.143' '0']
     ['5' '2.186' '0']
     ['6' '3.490' '0']]
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-03-08
      • 1970-01-01
      • 2013-03-13
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多