【发布时间】:2017-12-20 04:19:14
【问题描述】:
我正在尝试通过生成器将一维 numpy 数组(展平图像)输入 H5py 数据文件,以创建训练和验证矩阵。
以下代码改编自一个解决方案(现在找不到),其中 H5py 的 File 对象的 create_dataset 函数的 data 属性以调用 np.fromiter 的形式提供数据有一个生成器函数作为它的参数之一。
from scipy.misc import imread
import h5py
import numpy as np
import os
# Creating h5 data file
f = h5py.File('../data.h5', 'w')
# Source directory for image data
src = '/datasets/aic540/train/images/'
# Showing quantity and dimensionality of data
images = os.listdir(src)
ex_img = imread(src + images[0])
flat_img = ex_img.flatten()
print "# of images is {}".format(len(images))
print "image shape is {}".format(ex_img.shape)
print "flattened image shape is {}".format(flat_img.shape)
# Creating generator to feed in data to h5py's `create_dataset` function
gen = (imread(src + i).flatten().astype(np.int8) for i in os.listdir(src))
# Creating h5 dataset
f.create_dataset(name='training',
#shape=(59482, 1555200),
data=np.fromiter(gen, dtype=np.int8))
输出:
# of images is 59482
image shape is (540, 960, 3)
flattened image shape is (1555200,)
Traceback (most recent call last):
File "process_images.py", line 30, in <module>
data=np.fromiter(gen, dtype=np.int8))
ValueError: setting an array element with a sequence.
我在此上下文中搜索此错误时读到,问题在于np.fromiter() 需要一个列表而不是生成器函数(这似乎与名称“fromiter”所暗示的函数相反)——包装列表调用 list(gen) 中的生成器允许代码运行,但它当然会在调用 create_dataset 之前耗尽此列表扩展中的所有内存。
如何使用生成器将数据输入 H5py 数据文件?
如果我的方法完全错误,那么构建一个不适合内存的非常大的 numpy 矩阵的正确方法是什么——使用 H5py 还是其他方式?
【问题讨论】:
-
你必须写块。
np.fromiter(..., dtype=np.int8)创建一个数组 - 1d。因此,即使它可以从生成器生成数组,它仍然会在将整个内容传递给文件之前在内存中创建整个内容。 -
@hpaulj 与 ali_m 在这篇文章中建议的方式如此相似? stackoverflow.com/questions/34531479/… 看起来很不优雅/令人费解......我尝试使用
create_dataset函数的看起来更简单的chunk属性,但不幸的是这不起作用。