【问题标题】:NSMutableDictionary for huge dataset of floatsNSMutableDictionary 用于庞大的浮点数据集
【发布时间】:2012-01-20 03:32:52
【问题描述】:

我有一些代码可以将大型(数 GB)XML 文件转换为另一种格式。

除此之外,我需要在哈希表中存储 1 或 2 GB 的浮点数(每个条目有两个浮点数),并将 int 作为值的键。

目前,我正在使用 NSMutableDictionary 和一个包含两个浮点数的自定义类:

// create the dictionary
NSMutableDictionary *points = [[NSMutableDictionary alloc] init];

// add an entry (the data is read from an XML file using libxml)
int pointId = 213453;
float x = 42.313554; 
float y = -21.135213; 

MyPoint *point = [[MyPoint alloc] initWithX:x Y:y];
[points setObject:point forKey:[NSNumber numberWithInt:pointId]];
[point release];

// retrieve an entry (this happens later on while parsing the same XML file)
int pointId = 213453;
float x;
float y;
MyPoint *point = [points objectForKey:[NSNumber numberWithInt:pointId]];
x = point.x;
y = point.y;

这个数据集占用了我现在正在使用的 XML 文件大约 800MB 的 RAM,而且执行起来需要相当长的时间。我希望有更好的性能,但更重要的是我需要降低内存消耗,以便处理更大的 XML 文件。

objc_msg_send 就在代码配置文件中,- [NSNumber numberWithInt:] 也是如此,我确信我可以通过完全避免对象来降低内存使用量,但我对 C 编程了解不多(这个项目肯定在教我!)。

如何用高效的 C 数据结构替换 NSMuableDictionary、NSNumber MyPoint?没有任何第三方库依赖?

我还希望能够将这个数据结构写入磁盘上的文件,这样我就可以处理一个不完全适合内存的数据集,但我可能没有这个能力也可以生活。

(对于不熟悉 Objective-C 的人,NSMutableDictionary 类只能存储 Obj-C 对象,并且键也必须是对象。NSNumber 和 MyPoint 是哑容器类,允许 NSMutableDictionary 使用 float和 int 值。)

编辑:

我已经尝试使用 CFMutableDictionary 来存储结构,根据 apple's sample code。当字典为空时,它的性能很好。但随着字典的增长,它变得越来越慢。大约 25% 通过解析文件(字典中约 400 万个项目)它开始突突,比文件中的早期慢两个数量级。

NSMutableDictionary 没有相同的性能问题。 Instruments 展示了很多应用哈希和比较字典键的活动(下面的intEqual() 方法)。比较一个 int 很快,所以经常执行它是非常错误的。

这是我创建字典的代码:

typedef struct {
  float lat;
  float lon;
} AGPrimitiveCoord;

void agPrimitveCoordRelease(CFAllocatorRef allocator, const void *ptr) {
    CFAllocatorDeallocate(allocator, (AGPrimitiveCoord *)ptr);
}

Boolean agPrimitveCoordEqual(const void *ptr1, const void *ptr2) {
    AGPrimitiveCoord *p1 = (AGPrimitiveCoord *)ptr1;
    AGPrimitiveCoord *p2 = (AGPrimitiveCoord *)ptr2;

    return (fabsf(p1->lat - p2->lat) < 0.0000001 && fabsf(p1->lon - p2->lon) < 0.0000001);

}

Boolean intEqual(const void *ptr1, const void *ptr2) {
    return (int)ptr1 == (int)ptr2;
}

CFHashCode intHash(const void *ptr) {
  return (CFHashCode)((int)ptr);
}

// init storage dictionary
CFDictionaryKeyCallBacks intKeyCallBacks = {0, NULL, NULL, NULL, intEqual, intHash};
CFDictionaryValueCallBacks agPrimitveCoordValueCallBacks = {0, NULL /*agPrimitveCoordRetain*/, agPrimitveCoordRelease, NULL, agPrimitveCoordEqual};
temporaryNodeStore = CFDictionaryCreateMutable(NULL, 0, &intKeyCallBacks, &agPrimitveCoordValueCallBacks);

// add an item to the dictionary
- (void)parserRecordNode:(int)nodeId lat:(float)lat lon:(float)lon
{
  AGPrimitiveCoord *coordPtr = (AGPrimitiveCoord *)CFAllocatorAllocate(NULL, sizeof(AGPrimitiveCoord), 0);
  coordPtr->lat = lat;
  coordPtr->lon = lon;

  CFDictionarySetValue(temporaryNodeStore, (void *)nodeId, coordPtr);
}

编辑 2:

性能问题是由于 Apple 的示例代码中几乎无用的哈希实现造成的。我通过使用这个来提高性能:

// hash algorithm from http://burtleburtle.net/bob/hash/integer.html
uint32_t a = abs((int)ptr);
a = (a+0x7ed55d16) + (a<<12);
a = (a^0xc761c23c) ^ (a>>19);
a = (a+0x165667b1) + (a<<5);
a = (a+0xd3a2646c) ^ (a<<9);
a = (a+0xfd7046c5) + (a<<3);
a = (a^0xb55a4f09) ^ (a>>16);

【问题讨论】:

  • 部署后这个文件会改变吗?我正在考虑预处理。
  • 我永远不会在部署中使用此代码。这是一个命令行工具,它创建一个与我的应用程序一起部署的数据库(我从 openstreetmap.org 的 ~300GB 数据集中提取了几百兆字节)

标签: objective-c c performance hashmap nsmutabledictionary


【解决方案1】:

如果你想要类似 NSMutableDictionary 的行为但使用 malloc 的内存,你可以下拉到 CFDictionary(或者在你的情况下,CFMutableDictionary)。它实际上是 NSMutableDictionary 的基础,但它允许进行一些自定义,即您可以告诉它您没有存储对象。当您调用CFDictionaryCreateMutable() 时,您给它一个结构,描述您正在处理它的类型的值(它包含告诉它如何保留、释放、描述、散列和比较您的值的指针)。所以如果你想使用一个包含两个浮点数的结构,并且你很乐意为每个结构使用 malloc 的内存,你可以 malloc 结构,填充它,然后把它交给CFDictionary,然后你可以编写回调函数,以便它们与您的特定结构一起使用。您可以使用CFDictionary 的键和对象的唯一限制是它们需要适合void *。

【讨论】:

  • 我相信使用CFDictionaryCreate 可能会更快,因为该方法一次接收所有键和值,因此内部实现可能只需为所有记录分配一次内存。
  • @tia:如果您永远不必修改您的字典,那么请继续使用CFDictionaryCreate()。创建速度会更快(因为你说它不会随着时间的推移而增长字典),并且访问速度可能会更快(例如,可变版本可能使用效率较低的结构 -引擎盖,但这不一定是真的,无论如何这将是一个实现细节)。
  • 所以我正在关注stackoverflow.com/questions/1203726/… 中的示例代码,性能无法接受。当字典几乎为空时,它每秒添加大约 65000 个条目。当字典包含 400 万行时,它每秒只添加 600 个项目。 NSMutableDictionary 比这要快很多 很多,知道我做错了什么吗?我几乎完全使用了上述问题已接受答案中链接的 Apple 示例代码,并将编辑我的问题以包含我的代码。
  • @tia CFDictionaryCreate 将创建一个空字典。在解析 XML 文件时,我需要向其中添加 2000 万个项目,甚至在完成解析之前都不知道需要添加多少。所以 NS/CFMutableDictionary 是唯一的选择。
  • @AbhiBeckert:也许你需要一个更好的哈希函数?
【解决方案2】:

对于这类事情,我只会使用 C++ 容器 std::unordered_map 和 std::pair。您可以在 Objective-C++ 中使用它们。只需为您的文件添加 .mm 扩展名,而不是通常的 .m 扩展名。

更新

在您的评论中您说您以前从未使用过 C++。在这种情况下,您应该尝试 Kevin Ballard 对 CFDictionary 的回答,或者查看标准库中的 hcreate、hdestroy 和 hsearch 函数。

hcreate man page

【讨论】:

  • 你能给我一个例子来说明如何定义一个无序映射,添加一个条目,然后删除一个吗?我从来没有写过一行 C++。从数百兆数据集中随机访问值时,我应该期待什么类型的性能?
  • 谢谢。我查看了它,但无法理解手册页。 CFMutableDictionary 运行良好。
【解决方案3】:

将您的 .m 文件重命名为 .mm 并切换到使用 C++:

std::map<int, std::pair<float>> points;

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-08-24
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多