【问题标题】:Robomongo : Exceeded memory limit for $groupRobomongo :超过 $group 的内存限制
【发布时间】:2020-10-21 18:56:22
【问题描述】:

我正在使用脚本来删除 mongo 上的重复项,它在一个包含 10 个项目的集合中工作,我将其用作测试,但是当我用于包含 600 万个文档的真实集合时,出现错误。

这是我在 Robomongo(现在称为 Robo 3T)中运行的脚本:

var bulk = db.getCollection('RAW_COLLECTION').initializeOrderedBulkOp();
var count = 0;

db.getCollection('RAW_COLLECTION').aggregate([
  // Group on unique value storing _id values to array and count 
  { "$group": {
    "_id": { RegisterNumber: "$RegisterNumber", Region: "$Region" },
    "ids": { "$push": "$_id" },
    "count": { "$sum": 1 }      
  }},
  // Only return things that matched more than once. i.e a duplicate
  { "$match": { "count": { "$gt": 1 } } }
]).forEach(function(doc) {
  var keep = doc.ids.shift();     // takes the first _id from the array

  bulk.find({ "_id": { "$in": doc.ids }}).remove(); // remove all remaining _id matches
  count++;

  if ( count % 500 == 0 ) {  // only actually write per 500 operations
      bulk.execute();
      bulk = db.getCollection('RAW_COLLECTION').initializeOrderedBulkOp();  // re-init after execute
  }
});

// Clear any queued operations
if ( count % 500 != 0 )
    bulk.execute();

这是错误信息:

Error: command failed: {
    "errmsg" : "exception: Exceeded memory limit for $group, but didn't allow external sort. Pass allowDiskUse:true to opt in.",
    "code" : 16945,
    "ok" : 0
} : aggregate failed :
_getErrorWithCode@src/mongo/shell/utils.js:23:13
doassert@src/mongo/shell/assert.js:13:14
assert.commandWorked@src/mongo/shell/assert.js:266:5
DBCollection.prototype.aggregate@src/mongo/shell/collection.js:1215:5
@(shell):1:1

所以我需要设置allowDiskUse:true 才能工作?我在脚本中的哪个位置执行此操作,这样做有什么问题吗?

【问题讨论】:

    标签: mongodb duplicates out-of-memory


    【解决方案1】:
    { allowDiskUse: true } 
    

    应该放在聚合管道之后。

    在你的代码中应该是这样的:

    db.getCollection('RAW_COLLECTION').aggregate([
      // Group on unique value storing _id values to array and count 
      { "$group": {
        "_id": { RegisterNumber: "$RegisterNumber", Region: "$Region" },
        "ids": { "$push": "$_id" },
        "count": { "$sum": 1 }      
      }},
      // Only return things that matched more than once. i.e a duplicate
      { "$match": { "count": { "$gt": 1 } } }
    ], { allowDiskUse: true } )
    

    注意:使用{ allowDiskUse: true } 可能会引入与性能相关的问题,因为聚合管道将从磁盘上的临时文件中访问数据。还取决于磁盘性能和工作集的大小。测试用例的性能

    【讨论】:

    • 但是将其设置为 true 是否安全?我不明白为什么这是必要的
    • 聚合管道阶段有最大内存使用限制。要处理大型数据集,请将 allowDiskUse 选项设置为 true 以启用将数据写入临时文件。与完全从内存中读取相比,这应该提供不同的性能。还取决于数据集大小
    • 这给了我一个错误“AttributeError: 'dict' object has no attribute '_txn_read_preference'”。这是因为它不应该是一个dict。它应该是这样的:findings = list(collection.aggregate(aggr, allowDiskUse=True))跨度>
    【解决方案2】:

    当您拥有大量数据时,最好在分组之前使用匹配。 如果你在分组前使用匹配,你就不会遇到这个问题。

    db.getCollection('sample').aggregate([
       {$match:{State:'TAMIL NADU'}},
       {$group:{
           _id:{DiseCode:"$code", State:"$State"},
           totalCount:{$sum:1}
       }},
    
       {
         $project:{
            Code:"$_id.code",
            totalCount:"$totalCount",
            _id:0 
         }   
    
       }
    
    ])
    

    如果你真的在没有匹配的情况下克服了这个问题,那么解决方案是{ allowDiskUse: true }

    【讨论】:

      【解决方案3】:

      这是一个简单的无证技巧,可以在很多情况下帮助避免磁盘使用。

      您可以使用中间的$project 阶段来减小$sort 阶段中​​传递的记录的大小。

      在这个例子中,它将驱动到:

      var bulk = db.getCollection('RAW_COLLECTION').initializeOrderedBulkOp();
      var count = 0;
      
      db.getCollection('RAW_COLLECTION').aggregate([
        // here is the important stage
        { "$project": { "_id": 1, "RegisterNumber": 1, "Region": 1 } }, // this will reduce the records size
        { "$group": {
          "_id": { RegisterNumber: "$RegisterNumber", Region: "$Region" },
          "ids": { "$push": "$_id" },
          "count": { "$sum": 1 }      
        }},
        { "$match": { "count": { "$gt": 1 } } }
      ]).forEach(function(doc) {
        var keep = doc.ids.shift();     // takes the first _id from the array
      
        bulk.find({ "_id": { "$in": doc.ids }}).remove(); // remove all remaining _id matches
        count++;
      
        if ( count % 500 == 0 ) {  // only actually write per 500 operations
            bulk.execute();
            bulk = db.getCollection('RAW_COLLECTION').initializeOrderedBulkOp();  // re-init after execute
        }
      });
      

      查看第一个$project 阶段,这里只是为了避免磁盘使用。

      这对于收集大型记录特别有用,其中大部分数据未在聚合中使用

      【讨论】:

        【解决方案4】:

        From MongoDB Docs

        $group 阶段的 RAM 限制为 100 兆字节。默认情况下,如果 阶段超过此限制,$group 将产生错误。然而, 要允许处理大型数据集,请设置 allowDiskUse 选项为 true 以启用 $group 操作以写入临时 文件。请参阅 db.collection.aggregate() 方法和聚合命令 了解详情。

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 2012-11-11
          • 2016-08-24
          • 2017-04-04
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          相关资源
          最近更新 更多