【发布时间】:2015-01-08 20:33:53
【问题描述】:
我在 mongoshell 中有一个脚本,它应该从另一个(数据)中填充一个集合(数据聚合),每 5 分钟聚合一次时间序列。
数据收集有 7.000.000 多个条目,脚本需要很长时间才能完成...... 500.000 个数据需要 8 小时才能考虑在内,现在似乎已冻结。
基本上数据集合有这样的记录:
{
isodate: '2014-12-1OT12:47:32.000+02.00',
value: 234,
parentID: 123
}
数据聚合集合具有如下记录:
{
t: '2014-12-1OT12:45:00.000+02.00',
pid: 123, // parentID
sum: 1234, // sum of all the value of data between 12:45 and 12:50
count: 5, // number of data elements between 12:45 and 12:50
min: 23,
max: 435
}
数据集合的每条记录都将是数据聚合集合记录的一部分(在 count 属性中计为 1)。
// Cleanup collection
db.dataaggregation.remove({})
// Loop through data and populate the dataaggregation collection
db.data.find().addOption(DBQuery.Option.noTimeout).forEach(function(dt){
// Get 5 minutes timestamp
// eg: '2014-12-1OT12:47:32.000+02.00' => '2014-12-1OT12:45:00.000+02.00'
dt.isodate.setMinutes(dt.isodate.getMinutes() - dt.isodate.getMinutes() % 5);
dt.isodate.setSeconds(0);
// Create the dataaggregation record for the (timestamp, parentID) couple if does
// not exist or update the existing one
var d = db.dataaggregation.findOne({t: dt.isodate, pid: dt.parentID});
if(!d){
db.dataaggregation.insert({
t:dt.isodate,
pid: dt.parentID,
sum: dt.value,
count: 1,
min: dt.value,
max: dt.value
});
}else{
db.dataaggregation.update({
t:dt.isodate,
pid: dt.parentID
},{
$set:{
sum: d.sum + dt.value,
count: d.count + 1,
min: dt.value < d.min ? dt.value : d.min,
max: dt.value > d.max ? dt.value : d.max
}
},
{upsert:true}
);
}
})
有什么想法或建议来改进这一点吗?我有什么明显的遗漏吗?
【问题讨论】:
标签: mongodb mongodb-query aggregation-framework mongo-shell