【发布时间】:2016-07-21 08:30:56
【问题描述】:
鉴于 2 个巨大的值列表,我正在尝试使用 Scala 在 Spark 中计算它们之间的 jaccard similarity。
假设colHashed1 包含第一个值列表,colHashed2 包含第二个列表。
方法一(普通方法):
val jSimilarity = colHashed1.intersection(colHashed2).distinct.count/(colHashed1.union(colHashed2).distinct.count.toDouble)
方法2(使用minHashing):
我使用了here解释的方法。
import java.util.zip.CRC32
def getCRC32 (s : String) : Int =
{
val crc=new CRC32
crc.update(s.getBytes)
return crc.getValue.toInt & 0xffffffff
}
val maxShingleID = Math.pow(2,32)-1
def pickRandomCoeffs(kIn : Int) : Array[Int] =
{
var k = kIn
val randList = Array.fill(k){0}
while(k > 0)
{
// Get a random shingle ID.
var randIndex = (Math.random()*maxShingleID).toInt
// Ensure that each random number is unique.
while(randList.contains(randIndex))
{
randIndex = (Math.random()*maxShingleID).toInt
}
// Add the random number to the list.
k = k - 1
randList(k) = randIndex
}
return randList
}
val colHashed1 = list1Values.map(a => getCRC32(a))
val colHashed2 = list2Values.map(a => getCRC32(a))
val nextPrime = 4294967311L
val numHashes = 10
val coeffA = pickRandomCoeffs(numHashes)
val coeffB = pickRandomCoeffs(numHashes)
var signature1 = Array.fill(numHashes){0}
for (i <- 0 to numHashes-1)
{
// Evaluate the hash function.
val hashCodeRDD = colHashed1.map(ele => ((coeffA(i) * ele + coeffB(i)) % nextPrime))
// Track the lowest hash code seen.
signature1(i) = hashCodeRDD.min.toInt
}
var signature2 = Array.fill(numHashes){0}
for (i <- 0 to numHashes-1)
{
// Evaluate the hash function.
val hashCodeRDD = colHashed2.map(ele => ((coeffA(i) * ele + coeffB(i)) % nextPrime))
// Track the lowest hash code seen.
signature2(i) = hashCodeRDD.min.toInt
}
var count = 0
// Count the number of positions in the minhash signature which are equal.
for(k <- 0 to numHashes-1)
{
if(signature1(k) == signature2(k))
count = count + 1
}
val jSimilarity = count/numHashes.toDouble
方法 1 在时间方面似乎总是优于方法 2。当我分析代码时,在方法 2 中对 RDD 的 min() 函数调用需要大量时间,并且该函数被调用多次,具体取决于使用了多少哈希函数。
与重复的 min() 函数调用相比,方法 1 中使用的交集和并集操作似乎工作得更快。
我不明白为什么 minHashing 在这里没有帮助。与琐碎的方法相比,我希望 minHashing 工作得更快。我在这里做错了什么吗?
样本数据可以查看here
【问题讨论】:
-
您可以在数据集中为您的 col1 和 col2 添加示例数据吗?
-
@tuxdna 示例数据链接添加在问题的末尾
标签: scala apache-spark apache-spark-mllib