abhilb 的回答中肯且正确,但关于加载/读取文件还有很多话要说。快速的浏览器搜索将带您走很长一段路(我鼓励您阅读此内容!),但我将在此处添加一些详细信息:
如果你想加载多个匹配模式的文件,你可以通过 glob 迭代地这样做:
import pandas as pd
from glob import glob as gg
filePattern = "/path/to/file/*.txt"
for fileName in gg(filePattern):
df = pd.read_csv('OR0024622_auto3200.txt', delimiter=r'\t')
这将一个接一个地加载每个文件。如果要将所有数据放入单个数据框中怎么办?这样做:
masterDF = pd.Dataframe()
for fileName in gg(filePattern):
df = pd.read_csv('OR0024622_auto3200.txt', delimiter=r'\t')
masterDF = pd.concat([masterDF, df], axis=0)
这对 pandas 很有效,但是如果你想读入一个 numpy 数组呢?
import numpy as np
# using previous imports
base = 83
sensorlines = 6400
# create an empty array that has three columns
masterArray = np.full((0, 3), np.nan)
for fileName in gg(filePattern):
# open the file (NOTE: this does not read the file, just puts it in a buffer)
with open(fileName, "r") as tmp:
# now read the file and split each line by the carriage return (could be "\r\n")
# you now have a list of strings
data = tmp.read().split("\n")
# keep only the "data" portion of the file
data = data[base:sensorlines + base]
# convert list of strings to an array of floats
# here, I use a "list comprehension" for speed and simplicity
data = np.array([r.split("\t") for r in data]).astype(float)
# stack your new data onto your master array
masterArray = np.vstack([masterArray, data])
通过 "with open(fileName, "r")" 语法打开文件很方便,因为完成后 Python 会自动关闭文件。如果不使用“with”,则必须手动关闭文件(例如 tmp.close())。
这些只是帮助您上路的一些起点。随时要求澄清。