【问题标题】:Data from multiple sensors saved to txt file imported to pandas来自多个传感器的数据保存到导入 pandas 的 txt 文件
【发布时间】:2019-12-13 12:05:35
【问题描述】:

大家好。

我希望这里有人可以帮助我解决一些问题。我进行了一项实验,同时从 6 个独立的传感器收集数据。然后将数据导出到一个共享的 txt 文件中。现在我需要将数据导入python进行分析。

我知道我可以通过获取每一行并将每个传感器输出的数据复制并粘贴到一个单独的文档中来做到这一点,然后循环导入这些数据 - 但这是一项大量工作并且带来了很大的潜力人为错误。

但是有没有办法使用 readline 读取特定的行,并将其移植到 pandas DataFrame?每个传感器之间有固定的表头间距和行间距。

我试过了:

f=open('OR0024622_auto3200.txt')
lines = f.readlines()

base = 83
sensorlines = 6400

Sensor=[]
Sensor = lines[base:sensorlines+base]

df_sens = pd.DataFrame(Sensor)
df_sens

但输出不是很有用: Snip from of Output

-- 这是我要导入的文件: link.

有什么建议吗?

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    看起来像一个制表符分隔的数据。

    使用

    >>> df = pd.read_csv('OR0024622_auto3200.txt', delimiter=r'\t', skiprows=83, header=None, nrows=38955-84)
    >>> df.tail()
              0                   1              2
    38686  6397   3.1980000000e+003   9.28819e-009
    38687  6398   3.1985000000e+003   9.41507e-009
    38688  6399   3.1990000000e+003   1.11703e-008
    38689  6400   3.1995000000e+003   9.64276e-009
    38690  6401   3.2000000000e+003   8.92203e-009
    >>> df.head()
       0                   1              2
    0  1   0.0000000000e+000   6.62579e+000
    1  2   5.0000000000e-001   3.31289e+000
    2  3   1.0000000000e+000   2.62362e-011
    3  4   1.5000000000e+000   1.51130e-011
    4  5   2.0000000000e+000   8.35723e-012
    

    【讨论】:

    • 嘿 abhilb,是的,它是制表符分隔的 tata。问题真的是数量。查看我在 OP 中添加的链接。 :)
    • @SveinnGrétarsson 我向你保证这是一个小文件。有什么问题?
    • 数量是什么意思?
    • 感谢 abhilb!我确实需要添加 'engine='python'' 来避免解析错误,但这似乎很有魅力。
    【解决方案2】:

    abhilb 的回答中肯且正确,但关于加载/读取文件还有很多话要说。快速的浏览器搜索将带您走很长一段路(我鼓励您阅读此内容!),但我将在此处添加一些详细信息:

    如果你想加载多个匹配模式的文件,你可以通过 glob 迭代地这样做:

    import pandas as pd
    from glob import glob as gg
    filePattern = "/path/to/file/*.txt"
    
    for fileName in gg(filePattern):
        df = pd.read_csv('OR0024622_auto3200.txt', delimiter=r'\t')
    

    这将一个接一个地加载每个文件。如果要将所有数据放入单个数据框中怎么办?这样做:

    masterDF = pd.Dataframe()
    
    for fileName in gg(filePattern):
        df = pd.read_csv('OR0024622_auto3200.txt', delimiter=r'\t')
        masterDF = pd.concat([masterDF, df], axis=0)
    

    这对 pandas 很有效,但是如果你想读入一个 numpy 数组呢?

    import numpy as np
    
    # using previous imports
    base = 83
    sensorlines = 6400
    
    # create an empty array that has three columns    
    masterArray = np.full((0, 3), np.nan)
    
    for fileName in gg(filePattern):
        # open the file (NOTE: this does not read the file, just puts it in a buffer)
        with open(fileName, "r") as tmp:
            # now read the file and split each line by the carriage return (could be "\r\n")
            # you now have a list of strings
            data = tmp.read().split("\n")
    
            # keep only the "data" portion of the file
            data = data[base:sensorlines + base]
    
            # convert list of strings to an array of floats
            # here, I use a "list comprehension" for speed and simplicity
            data = np.array([r.split("\t") for r in data]).astype(float)
    
            # stack your new data onto your master array
            masterArray = np.vstack([masterArray, data])
    

    通过 "with open(fileName, "r")" 语法打开文件很方便,因为完成后 Python 会自动关闭文件。如果不使用“with”,则必须手动关闭文件(例如 tmp.close())。

    这些只是帮助您上路的一些起点。随时要求澄清。

    【讨论】:

    • 感谢 tnknepp。 :) 我会按照建议进行调查!
    猜你喜欢
    • 1970-01-01
    • 2020-05-28
    • 2013-02-03
    • 1970-01-01
    • 2018-03-01
    • 2016-07-08
    • 1970-01-01
    • 1970-01-01
    • 2020-03-24
    相关资源
    最近更新 更多