【问题标题】:Extracting data from text files using python [closed]使用python从文本文件中提取数据[关闭]
【发布时间】:2020-06-17 07:51:42
【问题描述】:

我有一个文本文件,其中包含这样的一行:

Component Sizing Information, AirTerminal:SingleDuct:VAV:Reheat, SPACE2-1 VAV REHEAT, Design Size Maximum Flow per Zone Floor Area during Reheat [m3/s-m2], 1.31927E-003

当数字之前的语句是(只是一个例子!)时,我想提取行尾的数字(1.31927E-003):

Design Size Maximum Flow per Zone Floor Area during Reheat [m3/s-m2]

事实上,文本文件中有几个关键语句,我需要分别提取紧随其后的数字。

你推荐什么库和方法? (使用python 3)。谢谢!

【问题讨论】:

    标签: python text-extraction


    【解决方案1】:

    重新模块

    Python 有一个正则表达式模块,该模块可用于从文本中进行基于编程模式的提取。

    re 是 Python 3 中的正则表达式模块。

    这是一种适用于您的特定情况的模式(但可能需要根据字符串的一致性进行更改)


    图案

    找出适合您的情况的模式 - 在您的情况下,我们可以确定以下内容:

    • 你有一个可以重复 0-9 的整数:

      `[0-9]+`
      
    • 你有一个小数点:

      `\.` # \ is used as an escape character for a literal . as . has a use in regex
      
    • 你有一串数字,它包含字母E和一个连字符-

      `[0-9E-]+`
      

    按顺序组合这些功能:

    pattern = r'[0-9]+\.[0-9E-]+'

    注意,在许多正则表达式示例中,字符串前的 r'...' 通常是一个 - r 表示一个原始字符串,可以更好地处理字符串中的潜在转义字符。


    Python 中的正则表达式

    我们需要将它编译为一个正则表达式(regex)对象: prog = re.compile(pattern)

    findall 方法将返回所有字符串(不重叠)的列表 - 还有其他方法,例如 re.search 和 re.match,它们具有其他特定输出:

    results = re.findall(prog, your_string)
    

    测试

    import re
    mystr = 'Component Sizin1..31927J-003ggnoor' \
            ' Ar1.31927E-003ea' \
            ' du' \
            'rin1g.31927E-003g Re' \
            'he1.t31927E-003at ' \
            '[m3/s-m1.34545457E-0032], 1.3' \
            '191.31927E-00327' \
            'E-01...31927E-00303'
    
    pattern = r'[0-9]+\.[0-9E-]+'
    prog = re.compile(pattern)
    results = re.findall(pattern, mystr)
    print(results)
    
    .........
    
    ['1.31927E-003', '1.34545457E-0032', '1.3191']
    

    要学习正则表达式需要练习(和良好的交互环境)——例如regex101

    【讨论】:

      【解决方案2】:

      如果你所有的行都相似,你可以拆分原始行并将数字提取为:

      string = "Component Sizing Information, AirTerminal:SingleDuct:VAV:Reheat, SPACE2-1 VAV REHEAT, Design Size Maximum Flow per Zone Floor Area during Reheat [m3/s-m2], 1.31927E-003"
      string = string.split(',')          #split the string at commas
      number = string[-1]                 #Extract the last number.
      number = number.strip()             #remove extra white spaces
      

      【讨论】:

        猜你喜欢
        • 2013-03-13
        • 2016-01-20
        • 2013-01-10
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2018-06-14
        • 2021-08-19
        • 1970-01-01
        相关资源
        最近更新 更多