【问题标题】:Java 8 Stream: groupingBy with multiple CollectorsJava 8 Stream:groupingBy 与多个收集器
【发布时间】:2015-08-18 11:54:56
【问题描述】:

我想使用 Java 8 Stream 并按一个分类器分组,但有多个收集器函数。因此,在分组时,例如计算一个字段(或另一个字段)的平均值和总和。

我试着用一个例子来简化一下:

public void test() {
    List<Person> persons = new ArrayList<>();
    persons.add(new Person("Person One", 1, 18));
    persons.add(new Person("Person Two", 1, 20));
    persons.add(new Person("Person Three", 1, 30));
    persons.add(new Person("Person Four", 2, 30));
    persons.add(new Person("Person Five", 2, 29));
    persons.add(new Person("Person Six", 3, 18));

    Map<Integer, Data> result = persons.stream().collect(
            groupingBy(person -> person.group, multiCollector)
    );
}

class Person {
    String name;
    int group;
    int age;

    // Contructor, getter and setter
}

class Data {
    long average;
    long sum;

    public Data(long average, long sum) {
        this.average = average;
        this.sum = sum;
    }

    // Getter and setter
}

结果应该是一个关联分组结果的Map

1 => Data(average(18, 20, 30), sum(18, 20, 30))
2 => Data(average(30, 29), sum(30, 29))
3 => ....

这对像“Collectors.counting()”这样的一个函数非常有效,但我喜欢链接多个(理想情况下从列表中无限)。

List<Collector<Person, ?, ?>>

有可能做这样的事情吗?

【问题讨论】:

  • 我是否理解正确,您的 Data 类只是平面图函数集合的占位符?这些函数中的每一个都必须对所有组执行相同的操作(例如,首先计算组的平均年龄,然后计算总年龄等)?
  • 当我说对了,我想是的。我只是用它在一个对象中保存多个数据,同时仍然能够识别哪个是哪个。也可以是 Key=FunctionName, Value=FunctionResult 的 Array 或 Map。

标签: java java-8 java-stream


【解决方案1】:

对于求和和平均的具体问题,使用collectingAndThen和summarizingDouble:

Map<Integer, Data> result = persons.stream().collect(
        groupingBy(Person::getGroup, 
                collectingAndThen(summarizingDouble(Person::getAge), 
                        dss -> new Data((long)dss.getAverage(), (long)dss.getSum()))));

对于更一般的问题(收集有关您的 Persons 的各种信息),您可以像这样创建一个复杂的收集器:

// Individual collectors are defined here
List<Collector<Person, ?, ?>> collectors = Arrays.asList(
        Collectors.averagingInt(Person::getAge),
        Collectors.summingInt(Person::getAge));

@SuppressWarnings("unchecked")
Collector<Person, List<Object>, List<Object>> complexCollector = Collector.of(
    () -> collectors.stream().map(Collector::supplier)
        .map(Supplier::get).collect(toList()),
    (list, e) -> IntStream.range(0, collectors.size()).forEach(
        i -> ((BiConsumer<Object, Person>) collectors.get(i).accumulator()).accept(list.get(i), e)),
    (l1, l2) -> {
        IntStream.range(0, collectors.size()).forEach(
            i -> l1.set(i, ((BinaryOperator<Object>) collectors.get(i).combiner()).apply(l1.get(i), l2.get(i))));
        return l1;
    },
    list -> {
        IntStream.range(0, collectors.size()).forEach(
            i -> list.set(i, ((Function<Object, Object>)collectors.get(i).finisher()).apply(list.get(i))));
        return list;
    });

Map<Integer, List<Object>> result = persons.stream().collect(
        groupingBy(Person::getGroup, complexCollector)); 

映射值是列表,其中第一个元素是应用第一个收集器的结果,依此类推。您可以使用 Collectors.collectingAndThen(complexCollector, list -&gt; ...) 添加自定义完成器步骤,以将此列表转换为更合适的内容。

【讨论】:

  • 好吧,这很有趣,我认为这与 Peter Lawrey 的回答有关。但这意味着它对我认为的每种功能都不灵活。我想使用函数列表(收集器)。使用您的解决方案,我认为我仅限于 summarizingDouble 所做的事情。
  • 哇!感谢你的回答。没想到会得到好东西。刚刚测试了它,结果就是我想要的。现在我只需要了解您的代码的所有内容。我认为这需要一些时间。比我可以稍微调整一下以完全匹配我试图完成的事情。但这确实是我想要的!
  • @PhilippS,它只是结合了各个收集器的各个操作(新供应商创建各个供应商的结果列表,新的累加器调用为每个单独的列表项累积等等)。如果您对此实现有任何具体问题,请随时询问。
  • 我会通过将所有内容放入工厂方法来重新设计。 collectors 可以是被 lambda 封闭的局部变量(实际上是方法参数)。完成器 lambda 可以另外应用用户提供的函数(在 OP 的情况下)将获取列表并返回 Data 对象。
  • 我提取了将给定收集器组合到一个方法中的代码,并在gist.github.com/dpolivaev/50cc9eb1b75453d37195882a9fc9fb69修复了泛型
【解决方案2】:

通过使用地图作为输出类型,可以有一个可能无限的 reducer 列表,每个减速器都会产生自己的统计数据并将其添加到地图中。

public static <K, V> Map<K, V> addMap(Map<K, V> map, K k, V v) {
    Map<K, V> mapout = new HashMap<K, V>();
    mapout.putAll(map);
    mapout.put(k, v);
    return mapout;
}

...

    List<Person> persons = new ArrayList<>();
    persons.add(new Person("Person One", 1, 18));
    persons.add(new Person("Person Two", 1, 20));
    persons.add(new Person("Person Three", 1, 30));
    persons.add(new Person("Person Four", 2, 30));
    persons.add(new Person("Person Five", 2, 29));
    persons.add(new Person("Person Six", 3, 18));

    List<BiFunction<Map<String, Integer>, Person, Map<String, Integer>>> listOfReducers = new ArrayList<>();

    listOfReducers.add((m, p) -> addMap(m, "Count", Optional.ofNullable(m.get("Count")).orElse(0) + 1));
    listOfReducers.add((m, p) -> addMap(m, "Sum", Optional.ofNullable(m.get("Sum")).orElse(0) + p.i1));

    BiFunction<Map<String, Integer>, Person, Map<String, Integer>> applyList
            = (mapin, p) -> {
                Map<String, Integer> mapout = mapin;
                for (BiFunction<Map<String, Integer>, Person, Map<String, Integer>> f : listOfReducers) {
                    mapout = f.apply(mapout, p);
                }
                return mapout;
            };
    BinaryOperator<Map<String, Integer>> combineMaps
            = (map1, map2) -> {
                Map<String, Integer> mapout = new HashMap<>();
                mapout.putAll(map1);
                mapout.putAll(map2);
                return mapout;
            };
    Map<String, Integer> map
            = persons
            .stream()
            .reduce(new HashMap<String, Integer>(),
                    applyList, combineMaps);
    System.out.println("map = " + map);

生产:

map = {Sum=10, Count=6}

【讨论】:

    【解决方案3】:

    您应该构建一个作为收集器聚合器的抽象,而不是链接收集器:使用一个接受收集器列表并将每个方法调用委托给每个收集器的类实现Collector 接口。然后,最后,您返回 new Data() 以及嵌套收集器产生的所有结果。

    您可以通过使用Collector.of(supplier, accumulator, combiner, finisher, Collector.Characteristics... characteristics) 来避免使用所有方法声明创建自定义类。finisher lambda 将调用每个嵌套收集器的终结器,然后返回 Data 实例。

    【讨论】:

    • 感谢您的回答。特别是结合其他答案,这提供了有关如何调整我的代码的良好背景知识。
    【解决方案4】:

    你可以把它们锁起来,

    一个收集器只能产生一个对象,但这个对象可以保存多个值。例如,您可以返回一个地图,其中地图为您返回的每个收集器都有一个条目。

    您可以使用Collectors.of(HashMap::new, accumulator, combiner);

    您的accumulator 将有一个收集器映射,其中生成的映射的键与收集器的名称匹配。当并行执行时,组合器需要一种组合多个结果的方法。


    通常,内置收集器使用数据类型来处理复杂结果。

    来自收藏家

    public static <T>
    Collector<T, ?, DoubleSummaryStatistics> summarizingDouble(ToDoubleFunction<? super T> mapper) {
        return new CollectorImpl<T, DoubleSummaryStatistics, DoubleSummaryStatistics>(
                DoubleSummaryStatistics::new,
                (r, t) -> r.accept(mapper.applyAsDouble(t)),
                (l, r) -> { l.combine(r); return l; }, CH_ID);
    }
    

    在它自己的类中

    public class DoubleSummaryStatistics implements DoubleConsumer {
        private long count;
        private double sum;
        private double sumCompensation; // Low order bits of sum
        private double simpleSum; // Used to compute right sum for non-finite inputs
        private double min = Double.POSITIVE_INFINITY;
        private double max = Double.NEGATIVE_INFINITY;
    

    【讨论】:

    • 我试图弄清楚这对我有什么帮助。也许我已经不在正确的轨道上了。这不只是一个像提到的“Collectors.counting()”这样的内置函数吗?但我想要做的是,使用其中两个或更多这些内置函数(在编译时未知)。或许你可以多解释一下。
    • 感谢您的回答。这提供了很好的背景知识,特别是理解 Tagirs 的答案并适应它。
    【解决方案5】:

    在 Java12 中,收集器 API 已扩展为静态 teeing(...) 函数:

    teeing​(Collector super T,​?,​R1> 下游1, 收藏家 super T,​?,​R2> 下游2, 双功能 合并)

    这提供了一种内置功能,可以在一个 Stream 上使用两个收集器并将结果合并到一个对象中。

    下面是一个小示例,其中将员工列表分成年龄组,对于每个组,根据年龄和薪水执行的两个 Collectors.summarizingInt() 作为IntSummaryStatistics 列表返回:

    import java.util.*;
    import java.util.function.Function;
    import java.util.stream.Collectors;
    
    public class CollectorTeeingTest {
    
    public static void main(String... args){
    
        NavigableSet<Integer> age_groups = new TreeSet<>();
        age_groups.addAll(List.of(30,40,50,60,Integer.MAX_VALUE)); //we don't want to map to null
    
        Function<Integer,Integer> to_age_groups = age -> age_groups.higher(age);
    
        List<Employee> employees = List.of( new Employee("A",21,2000),
                                            new Employee("B",24,2400),
                                            new Employee("C",32,3000),
                                            new Employee("D",40,4000),
                                            new Employee("E",41,4100),
                                            new Employee("F",61,6100)
        );
    
        Map<Integer,List<IntSummaryStatistics>> stats = employees.stream()
                .collect(Collectors.groupingBy(
                    employee -> to_age_groups.apply(employee.getAge()),
                    Collectors.teeing(
                        Collectors.summarizingInt(Employee::getAge),
                        Collectors.summarizingInt(Employee::getSalary),
                        (stat1, stat2) -> List.of(stat1,stat2))));
    
        stats.entrySet().stream().forEach(entry -> {
            System.out.println("Age-group: <"+entry.getKey()+"\n"+entry.getValue());
        });
    }
    
    public static class Employee{
    
        private final String name;
        private final int age;
        private final int salary;
    
        public Employee(String name, int age, int salary){
            
            this.name = name;
            this.age = age;
            this.salary = salary;
        }
        public String getName(){return this.name;}
        public int getAge(){return this.age;}
        public int getSalary(){return this.salary;}
    }
    

    }

    输出:

    Age-group: <2147483647
    [IntSummaryStatistics{count=1, sum=61, min=61, average=61,000000, max=61}, IntSummaryStatistics{count=1, sum=6100, min=6100, average=6100,000000, max=6100}]
    Age-group: <50
    [IntSummaryStatistics{count=2, sum=81, min=40, average=40,500000, max=41}, IntSummaryStatistics{count=2, sum=8100, min=4000, average=4050,000000, max=4100}]
    Age-group: <40
    [IntSummaryStatistics{count=1, sum=32, min=32, average=32,000000, max=32}, IntSummaryStatistics{count=1, sum=3000, min=3000, average=3000,000000, max=3000}]
    Age-group: <30
    [IntSummaryStatistics{count=2, sum=45, min=21, average=22,500000, max=24}, IntSummaryStatistics{count=2, sum=4400, min=2000, average=2200,000000, max=2400}]
    

    【讨论】:

      猜你喜欢
      • 2015-02-24
      • 2014-06-11
      • 2014-04-29
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多