【问题标题】:Handling large (object) datasets with PHP使用 PHP 处理大型(对象)数据集
【发布时间】:2010-03-29 18:17:08
【问题描述】:

我目前正在开展一个广泛依赖 EAV 模型的项目。两个实体作为它们的属性分别由一个模型表示,有时会扩展其他模型(或至少是基本模型)。

到目前为止,这已经很好地工作了,因为应用程序的大部分区域只依赖于过滤的实体集,而不是整个数据集。

然而,现在我需要解析整个数据集(即:所有实体及其所有属性),以便提供基于属性的排序/过滤算法。

该应用程序当前包含大约 2200 个实体,每个实体具有大约 100 个属性。每个实体都由一个模型(例如Client_Model_Entity)表示,并有一个名为$_attributes 的受保护属性,它是Attribute 对象的数组。

每个实体对象大约 500KB,这导致服务器上的负载令人难以置信。对于 2000 个实体,这意味着单个任务需要 1GB 的 RAM(和大量的 CPU 时间)才能工作,这是不可接受的。

是否有任何模式或通用方法来迭代如此大的数据集?分页并不是一个真正的选择,因为为了提供排序算法,必须考虑所有因素。

编辑:希望使事情更清晰的代码示例:

// code from the resource model
for ($i=0,$n=count($rowset);$i<$n;++$i)
{
    $clientEntity = new Client_Model_Entity($rowset[$i]);
    // getattributes gets all possible attributes from the db and creates models for them
    // this is actually the big resource hog, as one client can have 100 attributes
    $clientEntity->getAttributes(); 
    $this->_rows[$i] = $clientEntity;
    // memory usage has now increased by 500KB
    echo $i . ' : ' . memory_get_usage() . '<br />';
}

【问题讨论】:

    标签: php sorting data-structures iteration


    【解决方案1】:

    如果属性之间有很多共同点,您可以查看享元模式:http://en.wikipedia.org/wiki/Flyweight_pattern。这可能会显着减少表示模型所需的对象数量。

    【讨论】:

      【解决方案2】:

      一种解决方案可能是实现Iterator interface 并同时解析一个对象。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2023-04-02
        • 1970-01-01
        • 1970-01-01
        • 2012-09-22
        • 2022-01-06
        • 2017-09-24
        相关资源
        最近更新 更多