【问题标题】:Two Seq Comparison in ScalaScala中的两个序列比较
【发布时间】:2016-05-14 03:20:29
【问题描述】:

这是一个在数据更新中确定更新数据的场景。案例类对象的 Seq 具有它们的 ID,并且案例类具有覆盖 equals 方法,该方法对其属性进行数据比较,不包括 ID 字段。要找出任何更新的数据条目,我需要从数据库中检索数据。并且需要比较两个序列。什么是 Scala 方法来找出任何更新的对象?

只是有一个想法:创建一个以对象 ID 为键、对象为其值的映射。这可能行得通。

(更新)这是我出来的解决方案

val existingDataList: Foo = ...
val existingDataMap: Map[Long, Foo] = existingDataList.map(d => d.id -> d)(collection.breakOut)

// To find out updated data    
val updatedData = inputData.filter(d => existingDataMap.get(d.id) != d)

【问题讨论】:

    标签: scala


    【解决方案1】:

    如果我理解正确,您已经通过覆盖 equals 完成了大部分艰苦的工作 - 当然,您必须也相应地覆盖 hashCode,例如:

    case class Thing(id:Long, foo:String, bar:Int) {
      override def equals(that:Any):Boolean = {
        that match {
          case Thing(_, f, b) => (f == this.foo) && (b == this.bar)
          case _ => false
        }
      }
    
      override def hashCode:Int = {
        // A quick hack. You probably shouldn't do this for real; 
        // set id to 0 and use the default hashing algorithm:
        ScalaRunTime._hashCode(this.copy(id = 0))
      }
    }
    

    现在我们定义几个Thing 实例:

    val t1 = Thing(111, "t1", 1)
    val t1c = Thing(112, "t1", 1) // Same as t1 but with a new ID
    val t2 = Thing(222, "t2", 2)
    val t3 = Thing(333, "t3", 3)
    val t4 = Thing(444, "t4", 4)
    val t4m = Thing(444, "t4m", 4)  // Same as t4 but with a modified "foo"
    

    让我们制作几个序列:

    val orig = Seq(t1, t2, t3, t4)
    val mod = Seq(t1c, t2, t3, t4m)
    

    现在diff 告诉我们我们需要知道的一切:

    mod.diff(orig)
    // => returns Seq(t4m) - just what we wanted
    

    【讨论】:

    • 那行得通。是的,我还写了 hashCode 方法。
    【解决方案2】:

    所以,您有两个集合,并且您想在其中找到成对的对象,它们具有相同的 id,但数据不同,对吧? diff 并不是你真正想要的。

    这样就可以了:

    (first ++ second)
      .groupBy (_.id)
      .mapValues (_.toSet)
      .filterNot { case (_, v) => v.size != 2 }
      .values
      .map { v => v.head -> v.last }
    

    它会给你一个像 (first, second) 这样的元组列表,其中两个元素具有相同的 id,但数据不同。

    这假设您的 id 在每个集合中都是唯一的,并且每个 id 都出现在两个集合中。

    或者,如果您可以保证集合的大小相同,并且包含完全相同的 ID 集,您可以执行以下操作,虽然效率较低,但更简单:

         first.sortBy(_.id)
           .zip(second.sortBy(_.id))
           .filterNot { case (a, b) => a == b }
    

    【讨论】:

    • 感谢您展示几个高级 Scale 方法。
    猜你喜欢
    • 2016-08-24
    • 2017-10-24
    • 1970-01-01
    • 1970-01-01
    • 2021-04-26
    • 1970-01-01
    • 2015-07-19
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多