MongoDB中强大的统计框架Aggregation使用实例解析


当前第2页 返回上一页

首先上 $match, 取出上海学生

{$match:{'province':'上海'}}

接下来 用 $group 统计平均年龄

{$group:{_id:'$province',$avg:'$age'}}

$avg 是 $group的子命令,用于求平均值,类似的还有 $sum, $max ....
上面两个命令等价于

select province, avg(age) 
 from student 
 where province = '上海'
 group by province

下面是Java代码

Mongo m = new Mongo("localhost", 27017);
 DB db = m.getDB("test");
 DBCollection coll = db.getCollection("student");
 
 /*创建 $match, 作用相当于query*/
 DBObject match = new BasicDBObject("$match", new BasicDBObject("province", "上海"));
 
 /* Group操作*/
 DBObject groupFields = new BasicDBObject("_id", "$province");
 groupFields.put("AvgAge", new BasicDBObject("$avg", "$age"));
 DBObject group = new BasicDBObject("$group", groupFields);
 
 /* 查看Group结果 */
 AggregationOutput output = coll.aggregate(match, group); // 执行 aggregation命令
 System.out.println(output.getCommandResult());

输出结果:

{ "serverUsed" : "localhost/127.0.0.1:27017" ,    
 "result" : [ 
  { "_id" : "上海" , "AvgAge" : 32.09375}
  ] ,     
  "ok" : 1.0
 }

如此工程就结束了,再看另外一个需求

统计每个省各科平均成绩

首先更具数据库文档结构,subjects是数组形式,需要先‘劈'开,然后再进行统计

主要处理步骤如下:

1. 先用$unwind 拆数组 2. 按照 province, subject 分租并求各科目平均分

$unwind 拆数组

{$unwind:'$subjects'}

按照 province, subject 分组,并求平均分

{$group:{
   _id:{
     subjname:”$subjects.name”,  // 指定group字段之一 subjects.name, 并重命名为 subjname
     province:'$province'     // 指定group字段之一 province, 并重命名为 province(没变)
   },
   AvgScore:{
    $avg:”$subjects.score”    // 对 subjects.score 求平均
   }
 }

java代码如下:

Mongo m = new Mongo("localhost", 27017);
 DB db = m.getDB("test");
 DBCollection coll = db.getCollection("student");
 
 /* 创建 $unwind 操作, 用于切分数组*/
 DBObject unwind = new BasicDBObject("$unwind", "$subjects");
 
 /* Group操作*/
 DBObject groupFields = new BasicDBObject("_id", new BasicDBObject("subjname", "$subjects.name").append("province", "$province"));
 groupFields.put("AvgScore", new BasicDBObject("$avg", "$subjects.scores"));
 DBObject group = new BasicDBObject("$group", groupFields);
 
 /* 查看Group结果 */
 AggregationOutput output = coll.aggregate(unwind, group); // 执行 aggregation命令
 System.out.println(output.getCommandResult());

输出结果

{ "serverUsed" : "localhost/127.0.0.1:27017" , 
  "result" : [ 
   { "_id" : { "subjname" : "英语" , "province" : "海南"} , "AvgScore" : 58.1} , 
   { "_id" : { "subjname" : "数学" , "province" : "海南"} , "AvgScore" : 60.485} ,
   { "_id" : { "subjname" : "语文" , "province" : "江西"} , "AvgScore" : 55.538} , 
   { "_id" : { "subjname" : "英语" , "province" : "上海"} , "AvgScore" : 57.65625} , 
   { "_id" : { "subjname" : "数学" , "province" : "广东"} , "AvgScore" : 56.690} , 
   { "_id" : { "subjname" : "数学" , "province" : "上海"} , "AvgScore" : 55.671875} ,
   { "_id" : { "subjname" : "语文" , "province" : "上海"} , "AvgScore" : 56.734375} , 
   { "_id" : { "subjname" : "英语" , "province" : "云南"} , "AvgScore" : 55.7301 } ,
   .
   .
   .
   .
   "ok" : 1.0
 }

统计就此结束.... 稍等,似乎有点太粗糙了,虽然统计出来的,但是根本没法看,同一个省份的科目都不在一起。囧

接下来进行下加强,

支线任务: 将同一省份的科目成绩统计到一起( 即,期望 'province':'xxxxx', avgscores:[ {'xxx':xxx}, ....] 这样的形式)

要做的有一件事,在前面的统计结果的基础上,先用 $project 将平均分和成绩揉到一起,即形如下面的样子

{ "subjinfo" : { "subjname" : "英语" ,"AvgScores" : 58.1 } ,"province" : "海南" }

再按省份group,将各科目的平均分push到一块,命令如下:

$project 重构group结果

{$project:{province:"$_id.province", subjinfo:{"subjname":"$_id.subjname", "avgscore":"$AvgScore"}}

$使用 group 再次分组

{$group:{_id:"$province", avginfo:{$push:"$subjinfo"}}}

java 代码如下:

Mongo m = new Mongo("localhost", 27017);
DB db = m.getDB("test");
DBCollection coll = db.getCollection("student");
       
/* 创建 $unwind 操作, 用于切分数组*/
DBObject unwind = new BasicDBObject("$unwind", "$subjects");
       
/* Group操作*/
DBObject groupFields = new BasicDBObject("_id", new BasicDBObject("subjname", "$subjects.name").append("province", "$province"));
groupFields.put("AvgScore", new BasicDBObject("$avg", "$subjects.scores"));
DBObject group = new BasicDBObject("$group", groupFields);
       
/* Reshape Group Result*/
DBObject projectFields = new BasicDBObject();
projectFields.put("province", "$_id.province");
projectFields.put("subjinfo", new BasicDBObject("subjname","$_id.subjname").append("avgscore", "$AvgScore"));
DBObject project = new BasicDBObject("$project", projectFields);
       
/* 将结果push到一起*/
DBObject groupAgainFields = new BasicDBObject("_id", "$province");
groupAgainFields.put("avginfo", new BasicDBObject("$push", "$subjinfo"));
DBObject reshapeGroup = new BasicDBObject("$group", groupAgainFields);
 
/* 查看Group结果 */
AggregationOutput output = coll.aggregate(unwind, group, project, reshapeGroup);
System.out.println(output.getCommandResult());

结果如下:

{ "serverUsed" : "localhost/127.0.0.1:27017" , 
 "result" : [ 
    { "_id" : "辽宁" , "avginfo" : [ { "subjname" : "数学" , "avgscore" : 56.46666666666667} , { "subjname" : "英语" , "avgscore" : 52.093333333333334} , { "subjname" : "语文" , "avgscore" : 50.53333333333333}]} , 
    { "_id" : "四川" , "avginfo" : [ { "subjname" : "数学" , "avgscore" : 52.72727272727273} , { "subjname" : "英语" , "avgscore" : 55.90909090909091} , { "subjname" : "语文" , "avgscore" : 57.59090909090909}]} , 
    { "_id" : "重庆" , "avginfo" : [ { "subjname" : "语文" , "avgscore" : 56.077922077922075} , { "subjname" : "英语" , "avgscore" : 54.84415584415584} , { "subjname" : "数学" , "avgscore" : 55.33766233766234}]} , 
    { "_id" : "安徽" , "avginfo" : [ { "subjname" : "英语" , "avgscore" : 55.458333333333336} , { "subjname" : "数学" , "avgscore" : 54.47222222222222} , { "subjname" : "语文" , "avgscore" : 52.80555555555556}]} 
  .
  .
  .
  ] , "ok" : 1.0}


打赏

取消

感谢您的支持,我会继续努力的!

扫码支持
扫码打赏,您说多少就多少

打开支付宝扫一扫,即可进行扫码打赏哦

分享从这里开始,精彩与您同在

评论

管理员已关闭评论功能...