【问题标题】:R: How to identify and label cluster groups in a dendrogram (created by hclust)?R:如何在树状图中识别和标记集群组(由 hclust 创建)?
【发布时间】:2019-02-13 14:21:32
【问题描述】:

我使用 hclust 来识别数据中的集群,并确定这些集群的性质。以下是一个非常简化的版本:

gg <- c(1,2,4,3,3,15,16)
hh <- c(1,10,3,10,10,18,16)
z <- data.frame(gg,hh)
means <- apply(z,2,mean)
sds <- apply(z,2,sd)
nor <- scale(z,center=means,scale=sds) 
d <- dist(nor, method = "euclidean")
fit <- hclust(d, method="ward.D2")
plot(fit)
rect.hclust(fit, k=3, border="red")  
groups <- cutree(fit, k=3) 
aggregate(nor,list(groups),mean)

使用聚合我可以看到这三个集群包括一个在 gg 和 hh 变量上都具有低值的集群,一个具有低 gg 和平均 hh 的集群,以及一个具有高 gg 和高 hh 值的集群

如何查看这些在树状图上的位置(到目前为止,我只能通过检查组的大小并将它们与树状图上的大小进行比较来判断)?以及如何以某种方式在树状图中标记这些集群组(例如,在每个集群上添加类似“低”、“中”、“高”的名称)?我更喜欢基本 R 中的答案

【问题讨论】:

  • 你查看过dendextend中的color_branches吗?

标签: r dendrogram hclust


【解决方案1】:

不幸的是,如果不使用 dendextend 包,就没有可用于标记的简单选项。最接近的办法是使用rect.hclust() 公式中的border 参数为矩形着色......但这并不好玩。看看 - http://www.sthda.com/english/wiki/beautiful-dendrogram-visualizations-in-r-5-must-known-methods-unsupervised-machine-learning。

在这种有 2 列的情况下,我建议简单地绘制 z data.frame 并按您的 groups 直观地着色或分组。如果您标记这些点,这将进一步使其与树状图具有可比性。看这个例子:

# your data
gg <- c(1,2,4,3,3,15,16)
hh <- c(1,10,3,10,10,18,16)
z <- data.frame(gg,hh)

# a fun visualization function
visualize_clusters <- function(z, nclusters = 3, 
                           groupcolors = c("blue", "black", "red"), 
                           groupshapes = c(16,17,18), 
                           scaled_axes = TRUE){
  nor <- scale(z) # already defualts to use the datasets mean, sd)
  d <- dist(nor, method = "euclidean")
  fit <<- hclust(d, method = "ward.D2") # saves fit to the environment too
  groups <- cutree(fit, k = nclusters) 

  if(scaled_axes) z <- nor
  n <- nrow(z)
  plot(z, main = "Visualize Clusters",
       xlim = range(z[,1]), ylim = range(z[,2]),
       pch = groupshapes[groups], col = groupcolors[groups])
  grid(3,3, col = "darkgray") # dividing the plot into a grid of low, medium and high
  text(z[,1], z[,2], 1:n, pos = 4)

  centroids <- aggregate(z, list(groups), mean)[,-1]
  points(centroids, cex = 1, pch = 8, col = groupcolors)
  for(i in 1:nclusters){
    segments(rep(centroids[i,1],n), rep(centroids[i,2],n), 
             z[groups==i,1], z[groups==i,2], 
             col = groupcolors[i])
  }
  legend("topleft", bty = "n", legend = paste("Cluster", 1:nclusters), 
         text.col = groupcolors, cex = .8)
}

现在我们可以将它们绘制在一起:

par(mfrow = c(2,1))
visualize_clusters(z, nclusters = 3, groupcolors = c("blue", "black", "red"))
plot(fit); rect.hclust(fit, 3, border = rev(c("blue", "black", "red")))
par(mfrow = c(1,1)

记下您的眼科检查的低-低、低-中、高-高的网格。

我喜欢线段。尝试使用更大的数据,例如:

gg <- runif(30,1,20)
hh <- c(runif(10,5,10),runif(10,10,20),runif(10,1,5))
z <- data.frame(gg,hh)
visualize_clusters(z, nclusters = 3, groupcolors = c("blue", "black", "red"))

希望这会有所帮助。

【讨论】:

  • 谢谢,我实际上有 3 个变量用作集群变量,但我很欣赏你的图,并且可能会在未来的某个时候使用它们(我用简单的散点图比较两个变量时间按组着色,但我喜欢你如何说明每个组的中心点)。我决定根据您的建议更改 rect.hclust 的边框颜色并添加一个图例。我想我真的不需要让 R 将组映射到树状图 - 因为我可以从组大小中弄清楚 - 我只是觉得这很酷。
猜你喜欢
  • 2011-01-19
  • 2015-11-05
  • 1970-01-01
  • 1970-01-01
  • 2019-01-07
  • 2020-07-19
  • 2015-12-15
  • 1970-01-01
  • 2016-04-29
相关资源
最近更新 更多