【问题标题】:How to trouble-shoot HDFStore Exception: cannot find the correct atom type如何解决 HDFStore 异常:找不到正确的原子类型
【发布时间】:2013-03-07 11:55:34
【问题描述】:

我正在寻找一些关于哪些类型的数据场景会导致此异常的一般指导。我曾尝试以各种方式按摩我的数据,但均无济于事。

我已经用谷歌搜索这个异常好几天了,经历了几次谷歌小组讨论,并没有想出调试HDFStore Exception: cannot find the correct atom type 的解决方案。我正在阅读一个混合数据类型的简单 csv 文件:

Int64Index: 401125 entries, 0 to 401124
Data columns:
SalesID                     401125  non-null values
SalePrice                   401125  non-null values
MachineID                   401125  non-null values
ModelID                     401125  non-null values
datasource                  401125  non-null values
auctioneerID                380989  non-null values
YearMade                    401125  non-null values
MachineHoursCurrentMeter    142765  non-null values
UsageBand                   401125  non-null values
saledate                    401125  non-null values
fiModelDesc                 401125  non-null values
Enclosure_Type              401125  non-null values
...................................................
Stick_Length                401125  non-null values
Thumb                       401125  non-null values
Pattern_Changer             401125  non-null values
Grouser_Type                401125  non-null values
Backhoe_Mounting            401125  non-null values
Blade_Type                  401125  non-null values
Travel_Controls             401125  non-null values
Differential_Type           401125  non-null values
Steering_Controls           401125  non-null values
dtypes: float64(2), int64(6), object(45)

存储数据帧的代码:

In [30]: store = pd.HDFStore('test0.h5','w')
In [31]: for chunk in pd.read_csv('Train.csv', chunksize=10000):
   ....:     store.append('df', chunk, index=False)

请注意,如果我在一次导入的数据帧上使用store.put,我可以成功存储它,尽管速度很慢(我相信这是由于对象 dtypes 的酸洗,即使对象只是字符串数据)。

是否存在可能引发此异常的 NaN 值注意事项?

例外:

Exception: cannot find the correct atom type -> [dtype->object,items->Index([Usa
geBand, saledate, fiModelDesc, fiBaseModel, fiSecondaryDesc, fiModelSeries, fiMo
delDescriptor, ProductSize, fiProductClassDesc, state, ProductGroup, ProductGrou
pDesc, Drive_System, Enclosure, Forks, Pad_Type, Ride_Control, Stick, Transmissi
on, Turbocharged, Blade_Extension, Blade_Width, Enclosure_Type, Engine_Horsepowe
r, Hydraulics, Pushblock, Ripper, Scarifier, Tip_Control, Tire_Size, Coupler, Co
upler_System, Grouser_Tracks, Hydraulics_Flow, Track_Type, Undercarriage_Pad_Wid
th, Stick_Length, Thumb, Pattern_Changer, Grouser_Type, Backhoe_Mounting, Blade_
Type, Travel_Controls, Differential_Type, Steering_Controls], dtype=object)] lis
t index out of range

更新 1

Jeff 关于存储在数据框中的列表的提示让我研究了嵌入式逗号。 pandas.read_csv 正在正确解析文件,并且双引号中确实有一些嵌入的逗号。所以这些字段本身不是 python 列表,但在文本中确实有逗号。以下是一些示例:

3     Hydraulic Excavator, Track - 12.0 to 14.0 Metric Tons
6     Hydraulic Excavator, Track - 21.0 to 24.0 Metric Tons
8       Hydraulic Excavator, Track - 3.0 to 4.0 Metric Tons
11      Track Type Tractor, Dozer - 20.0 to 75.0 Horsepower
12    Hydraulic Excavator, Track - 19.0 to 21.0 Metric Tons

但是,当我从 pd.read_csv 块中删除此列并附加到我的 HDFStore 时,我仍然得到相同的异常。当我尝试单独附加每一列时,我得到以下新异常:

In [6]: for chunk in pd.read_csv('Train.csv', header=0, chunksize=50000):
   ...:     for col in chunk.columns:
   ...:         store.append(col, chunk[col], data_columns=True)

Exception: cannot properly create the storer for: [_TABLE_MAP] [group->/SalesID
(Group) '',value-><class 'pandas.core.series.Series'>,table->True,append->True,k
wargs->{'data_columns': True}]

我会继续排查问题。这是数百条记录的链接:

https://docs.google.com/spreadsheet/ccc?key=0AutqBaUiJLbPdHFvaWNEMk5hZ1NTNlVyUVduYTZTeEE&usp=sharing

更新 2

好的,我在工作电脑上试了下,结果如下:

In [4]: store = pd.HDFStore('test0.h5','w')

In [5]: for chunk in pd.read_csv('Train.csv', chunksize=10000):
   ...:     store.append('df', chunk, index=False, data_columns=True)
   ...:

Exception: cannot find the correct atom type -> [dtype->object,items->Index([fiB
aseModel], dtype=object)] [fiBaseModel] column has a min_itemsize of [13] but it
emsize [9] is required!

我想我知道这里发生了什么。如果我将字段 fiBaseModel 的最大长度作为第一个块,我会得到:

In [16]: lens = df.fiBaseModel.apply(lambda x: len(x))

In [17]: max(lens[:10000])
Out[17]: 9

还有第二块:

In [18]: max(lens[10001:20000])
Out[18]: 13

因此,为该列创建了 9 字节的存储表,因为这是第一个块的最大值。当它在后续的块中遇到较长的文本字段时,它会抛出异常。

我对此的建议是要么截断后续块中的数据(带有警告),要么允许用户指定列的最大存储空间并截断超过它的任何内容。也许 pandas 已经可以做到了,我还没有时间真正深入了解HDFStore。

更新 3

尝试使用 pd.read_csv 导入 csv 数据集。我将所有对象的字典传递给 dtypes 参数。然后我遍历文件并将每个块存储到 HDFStore 中,并为min_itemsize 传递一个大值。我得到以下异常:

AttributeError: 'NoneType' object has no attribute 'itemsize'

我的简单代码:

store = pd.HDFStore('test0.h5','w')
objects = dict((col,'object') for col in header)

for chunk in pd.read_csv('Train.csv', header=0, dtype=objects,
    chunksize=10000, na_filter=False):
    store.append('df', chunk, min_itemsize=200)

我已尝试调试并检查堆栈跟踪中的项目。这是异常时表格的样子:

ipdb> self.table
/df/table (Table(10000,)) ''
  description := {
  "index": Int64Col(shape=(), dflt=0, pos=0),
  "values_block_0": StringCol(itemsize=200, shape=(53,), dflt='', pos=1)}
  byteorder := 'little'
  chunkshape := (24,)
  autoIndex := True
  colindexes := {
    "index": Index(6, medium, shuffle, zlib(1)).is_CSI=False}

更新 4

现在我正在尝试迭代地确定我的数据框对象列中最长字符串的长度。我就是这样做的:

    def f(x):
        if x.dtype != 'object':
            return
        else:
            return len(max(x.fillna(''), key=lambda x: len(str(x))))

lengths = pd.DataFrame([chunk.apply(f) for chunk in pd.read_csv('Train.csv', chunksize=50000)])
lens = lengths.max().dropna().to_dict()

In [255]: lens
Out[255]:
{'Backhoe_Mounting': 19.0,
 'Blade_Extension': 19.0,
 'Blade_Type': 19.0,
 'Blade_Width': 19.0,
 'Coupler': 19.0,
 'Coupler_System': 19.0,
 'Differential_Type': 12.0
 ... etc... }

一旦我有了最大字符串列长度的字典,我会尝试通过min_itemsize 参数将它传递给append:

In [262]: for chunk in pd.read_csv('Train.csv', chunksize=50000, dtype=types):
   .....:     store.append('df', chunk, min_itemsize=lens)

Exception: cannot find the correct atom type -> [dtype->object,items->Index([Usa
geBand, saledate, fiModelDesc, fiBaseModel, fiSecondaryDesc, fiModelSeries, fiMo
delDescriptor, ProductSize, fiProductClassDesc, state, ProductGroup, ProductGrou
pDesc, Drive_System, Enclosure, Forks, Pad_Type, Ride_Control, Stick, Transmissi
on, Turbocharged, Blade_Extension, Blade_Width, Enclosure_Type, Engine_Horsepowe
r, Hydraulics, Pushblock, Ripper, Scarifier, Tip_Control, Tire_Size, Coupler, Co
upler_System, Grouser_Tracks, Hydraulics_Flow, Track_Type, Undercarriage_Pad_Wid
th, Stick_Length, Thumb, Pattern_Changer, Grouser_Type, Backhoe_Mounting, Blade_
Type, Travel_Controls, Differential_Type, Steering_Controls], dtype=object)] [va
lues_block_2] column has a min_itemsize of [64] but itemsize [58] is required!

违规列传递的 min_itemsize 为 64,但异常状态要求 itemsize 为 58。这可能是一个错误?

在 [266] 中:pd.版本 输出[266]:'0.11.0.dev-eb07c5a'

【问题讨论】:

  • nan 没问题,即使在字符串中也是如此。对象类型必须是纯字符串(而不是其他任何东西)。显示几行我的实际数据。
  • 通过尝试一次存储一个列来解决问题,直到您失败;错误消息是最后一个字段。传递 data_columns=True 来追加(实际上你不会这样做,因为你可能不需要查询所有列)
  • 这些列中除了字符串之外还有其他内容,也许是一个列表?
  • 这是无效的。你告诉它一个特定的列应该有一个最小长度,但是所有的字符串都在一个块中,因为你没有指定 data_columns (这是一个错误,我让你指定这个 - 错误消息应该更好),你如果它们是单独的字段,则每列只能有一个最小值,否则最小值是一个(每个都相同)块中的每列。只需使用 min 的 min (请注意,您实际上是在查看数据两次才能做到这一点)。最好只选择一个合理的数字(单独的问题是我们可以让你截断超出这个数字)
  • github.com/pydata/pandas/pull/3167 将在此处为您的代码提供错误(因为您正在尝试使用不可查询的列来 min_itemsize)。此外,您可以使用 lib.max_len_string_array(s.values) 在您的应用中快速获得最大值(更快,您不需要测试对象类型)

标签: python pandas hdf5


【解决方案1】:

您提供的链接可以很好地存储框架。逐列仅表示指定 data_columns=True。它将单独处理这些列并在有问题的列上提出。

诊断

store = pd.HDFStore('test0.h5','w')
In [31]: for chunk in pd.read_csv('Train.csv', chunksize=10000):
   ....:     store.append('df', chunk, index=False, data_columns=True)

在生产中,您可能希望将 data_columns 限制为要查询的列(也可以是 None,在这种情况下您只能查询索引/列)

更新:

您可能会遇到另一个问题。 read_csv 根据在每个块中看到的内容转换 dtypes, 所以对于 10,000 的块大小,附加操作失败,因为块 1 和 2 有 在某些列中查找整数数据,然后在块 3 中您有一些 NaN,所以它是因为浮动。 要么预先指定 dtypes,使用更大的块大小,要么运行你的操作两次 以保证您在块之间的 dtypes。

我已经更新了 pytables.py 在这种情况下有一个更有用的例外(以及 告诉您列是否包含不兼容的数据)

感谢您的报告!

【讨论】:

  • 变长字符数据的情况如何?
  • HDF5 在表格中不支持此功能(您可以通过非表格来实现,例如 put)。要么制作一个大列(这可以工作,但在空间方面可能效率低下),或者将这些数据分离出来并通过 put 存储(当然,当你查询时你必须处理它)。在此处查看建议:github.com/pydata/pandas/issues/3032
  • 您还可以查看 HDF5 部分(即将重命名为 HDFStore),此处为 pandas.pydata.org/pandas-docs/dev/cookbook.html
猜你喜欢
  • 2013-03-18
  • 2013-05-01
  • 1970-01-01
  • 1970-01-01
  • 2013-07-08
  • 2015-07-18
  • 2018-01-17
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多