能否将您的问题视为词袋模型,其中每篇文章(观察行)的词条不超过 100 个?
无论如何,我认为您必须提供有关“为什么”和“如何”对这些数据进行聚类的更多信息和示例。例如,我们有:
1 2 3
2 3 4
2 3 4 5
1 2 3 4
3 4 6
6 7 8
9 10
9 11
10 12 13 14
您期望的聚类是什么?这个聚类中有多少个聚类?只有两个集群?
在您提供更多信息之前,根据您目前的描述,我认为您不需要集群算法,而是需要连接组件的结构。第一轮处理数据集以获取连接组件的信息,您需要第二轮检查每一行属于哪些连接组件。以上面的例子,第一轮:
1 2 3 : 1 <- 1, 1 <- 2, 1 <- 3 (all point linked to the smallest point to
represent they are belong to the same cluster of the smallest point)
2 3 4 : 2 <- 4 (2 and 3 have already linked to 1 which is <= 2, so they do
not need to change)
2 3 4 5 : 2 <- 5
1 2 3 4 : 1 <- 4 (in fact this change are not essential because we have
1 <- 2 <- 4, but change this can speed up the second round)
3 4 6 : 3 <- 6
6 7 8 : 6 <- 7, 6 <- 8
9 10 : 9 <- 9, 9 <- 10
9 11 : 9 <- 11
10 11 12 13 14 : 10 <- 12, 10 <- 13, 10 <- 14
现在我们有一个森林结构来表示点的连通分量。第二轮你可以轻松地在每一行中拾取一个点(最小的就是最好的)并在森林中追踪它的根。用你的话来说,具有相同根的行在同一个簇中。例如:
1 2 3 : 1 <- 1, cluster root 1
2 3 4 5 : 1 <- 1 <- 2, cluster root 1
6 7 8 : 1 <- 1 <- 3 <- 6, cluster root 1
9 10 : 9 <- 9, cluster root 9
10 11 12 13 14 : 9 <- 9 <- 10, cluster root 9
这个过程需要 O(k) 空间,其中 k 是点数,以及 O(nm + nh) 时间,其中r 是森林结构的高度,其中 r m。
我不确定这是否是您想要的结果。