【问题标题】:Appending to h5 files附加到 h5 文件
【发布时间】:2021-05-08 04:32:30
【问题描述】:

我有一个 h5 文件,其中包含这样的数据集:

col1.      col2.      col3
 1           3          5
 5           4          9
 6           8          0
 7           2          5
 2           1          2

我有另一个具有相同列的 h5 文件:

col1.      col2.      col3
 6           1          9
 8           2          7

我想将这两者连接起来,得到以下 h5 文件:

col1.      col2.      col3
 1           3          5
 5           4          9
 6           8          0
 7           2          5
 2           1          2
 6           1          9
 8           2          7

如果文件很大或者我们有很多这样的合并,最有效的方法是什么?

【问题讨论】:

  • h5_1.append(h5_2)?
  • 它们是熊猫数据框吗?如果是这样h5_concat = pandas.concat(h5_1, h5_2)。及时,这不是合并。是串联
  • 它们不是熊猫数据框。它们是两个 h5 文件。
  • pd.concat([h5_1,h5_2], axis=0)
  • @wwnde 你是否建议先将 h5 文件转换为 pandas 数据帧?

标签: python hdf5 h5py


【解决方案1】:

我不熟悉 pandas,所以无能为力。这可以通过 h5py 或 pytables 来完成。正如@hpaulj 提到的,该过程将数据集读入一个 numpy 数组,然后使用 h5py 写入一个 HDF5 数据集。确切的过程取决于 maxshape 属性(它控制是否可以调整数据集的大小)。

我创建了示例来展示这两种方法(固定大小或可调整大小的数据集)。第一种方法创建一个新的 file3,它结合了 file1 和 file2 的值。第二种方法将 file2 中的值添加到 file1e(可调整大小)。注意:创建示例中使用的文件的代码在最后。

关于 SO,我有一个更长的答案,它显示了复制数据的所有方法。
看到这个答案:How can I combine multiple .h5 file?

方法一:将数据集合并到一个新文件中
未使用maxshape= 参数创建数据集时需要

with h5py.File('file1.h5','r') as h5f1,  \
     h5py.File('file2.h5','r') as h5f2,  \
     h5py.File('file3.h5','w') as h5f3 :
         
    print (h5f1['ds_1'].shape, h5f1['ds_1'].maxshape)
    print (h5f2['ds_2'].shape, h5f2['ds_2'].maxshape)    

    arr1_a0 = h5f1['ds_1'].shape[0]            
    arr2_a0 = h5f2['ds_2'].shape[0]            
    arr3_a0 = arr1_a0 + arr2_a0          
    h5f3.create_dataset('ds_3', dtype=h5f1['ds_1'].dtype,
                        shape=(arr3_a0,3), maxshape=(None,3))

    xfer_arr1 = h5f1['ds_1']               
    h5f3['ds_3'][0:arr1_a0, :] = xfer_arr1
 
    xfer_arr2 = h5f2['ds_2']   
    h5f3['ds_3'][arr1_a0:arr3_a0, :] = xfer_arr2

    print (h5f3['ds_3'].shape, h5f3['ds_3'].maxshape)

方法 2:将 file2 数据集附加到 file1 数据集
file1e 中的数据集必须使用maxshape= 参数创建

with h5py.File('file1e.h5','r+') as h5f1, \
     h5py.File('file2.h5','r') as h5f2 :

    print (h5f1['ds_1e'].shape, h5f1['ds_1e'].maxshape)
    print (h5f2['ds_2'].shape, h5f2['ds_2'].maxshape)    
    
    arr1_a0 = h5f1['ds_1e'].shape[0]            
    arr2_a0 = h5f2['ds_2'].shape[0] 
    arr3_a0 = arr1_a0 + arr2_a0          

    h5f1['ds_1e'].resize(arr3_a0,axis=0)
    
    xfer_arr2 = h5f2['ds_2']   
    h5f1['ds_1e'][arr1_a0:arr3_a0, :] = xfer_arr2

    print (h5f1['ds_1e'].shape, h5f1['ds_1e'].maxshape)

创建上述示例文件的代码:

import h5py
import numpy as np

arr1 = np.array([[ 1, 3, 5 ],
                 [ 5, 4, 9 ],
                 [ 6, 8, 0 ],
                 [ 7, 2, 5 ],
                 [ 2, 1, 2 ]] )

with h5py.File('file1.h5','w') as h5f:
    h5f.create_dataset('ds_1',data=arr1)
    print (h5f['ds_1'].maxshape)   
    
with h5py.File('file1e.h5','w') as h5f:
    h5f.create_dataset('ds_1e',data=arr1, shape=(5,3), maxshape=(None,3))
    print (h5f['ds_1e'].maxshape)             
                 
arr2 = np.array([[ 6, 1, 9 ],
                 [ 8, 2, 7 ]] )
                 
with h5py.File('file2.h5','w') as h5f:
    h5f.create_dataset('ds_2',data=arr2)

【讨论】:

  • h5 文件将数据存储在数据集中。 h5f1.keys() 在根级别生成对象名称列表。在您的情况下,它们是名为“col1”、“col2”、“col3”的数据集。 h5f2.keys() 是否产生相同的名称?如果是这样,您是否要将h5f2['col1 '] 到h5f1['col1 '] 的数据组合起来,'col2' 和'col3' 也一样?如果是这样,这对于 3 个数据集是相同的过程。我是否需要修改我的示例以显示如何通过键/数据集进行迭代?它会“稍微复杂一些”。
  • 感谢您的回答。请告诉我是否有任何方法可以将h5f2['col1 '] 直接附加到h5f1['col1 '],而不是创建一个新数据集为h5f3['col1'] 并将这两个顺序添加到其中?
  • 示例的第二部分就是这样做的。它以附加模式打开'file1e.h5':r+,调整数据集的大小,然后附加来自'file2.h5' 的数据。附加到数据集需要在最初创建时将其定义为“可调整大小”(使用示例中所示的maxshape= 参数)。 0 轴的值必须是:a) None,它允许无限大小,或 b) 大于 h5f1['col1 '] 和 h5f2['col1 '] 之和的值。您需要为文件中的所有 3 个数据集检查此属性。
  • 第二部分是我要找的。非常感谢您的帮助。
猜你喜欢
  • 2020-08-30
  • 2021-03-29
  • 1970-01-01
  • 2015-05-11
  • 2020-07-15
  • 2013-04-23
  • 2017-06-10
  • 2017-08-07
  • 1970-01-01
相关资源
最近更新 更多