【问题标题】:How to print the data of each cluster in agglomerative clustering algorithm in python如何在python中的凝聚聚类算法中打印每个聚类的数据
【发布时间】:2020-07-07 00:51:24
【问题描述】:

我是python机器学习工具的新手,我编写了这个凝聚层次聚类的代码,但我不知道是否有任何方法可以打印每个绘图集群的数据。 算法的输入是5个数字(0,1,2,3,4),除了绘制簇,我需要单独打印每个簇的值,如下所示 cluster1= [1,2,4] cluster2=[0,3]

更新:我想得到根据这条线和其他线plt.scatter(points[y_hc==0,0], points[y_hc==0,1],s=100,c='cyan')绘制和着色的数据,根据这段代码,这些数字(1,2,4)合二为一cluster 并且具有相同的颜色,并且 (0,3) 在 cluster2 中,因此,我需要在终端中打印这些数据(每个集群的数据)。这段代码只是绘制数据。

import numpy as np 
import matplotlib.pyplot as plt 
from sklearn.datasets import make_blobs
dataset= make_blobs(n_samples=5, n_features=2,centers=4, cluster_std=1.6, random_state=50)
points= dataset[0]

import scipy.cluster.hierarchy as sch 
from sklearn.cluster import AgglomerativeClustering

dendrogram = sch.dendrogram(sch.linkage(points,method='ward'))
plt.scatter(dataset[0][:,0],dataset[0][:,1])
hc = AgglomerativeClustering(n_clusters=4, affinity='euclidean',linkage='ward')
y_hc= hc.fit_predict(points)
plt.scatter(points[y_hc==0,0], points[y_hc==0,1],s=100,c='cyan')
plt.scatter(points[y_hc==1,0], points[y_hc==1,1],s=100,c='yellow')
plt.scatter(points[y_hc==2,0], points[y_hc==2,1],s=100,c='red')
plt.scatter(points[y_hc==3,0], points[y_hc==3,1],s=100,c='green')
plt.show()

【问题讨论】:

  • 不太清楚您要达到的目标。你遇到的问题是什么?使用您发布的代码,您不会看到所有的图,因为您最后只调用了一次plt.show()。当我运行您的代码时(顺便说一句,感谢您提供了一个可重现的示例!)我得到的散点图似乎与y_hc 中的预测正确对应。使用plt.scatter(dataset[0][:, 0], dataset[0][:, 1], c=y_hc); 可以更简洁地获得类似的结果。
  • 这段代码只是在树状图中对输入值(5 个数字)进行聚类,正如您在运行代码时看到的那样,因此,我需要在终端中打印引用每个聚类的值。
  • 在树状图中,值(0,3)是指集群,(2.1,4)是指其他集群,所以,我需要获取这个值并在终端打印
  • 好的,抱歉,我有点慢。您的意思是将dataset[1] 中的值与dendrogram 变量(dict) 中的集群相匹配吗?
  • 我需要将树状图中显示的每个集群的值保存在数组或任何内容中,然后打印出来

标签: python algorithm machine-learning hierarchical-clustering


【解决方案1】:

经过一些研究,似乎没有一种简单的方法可以从scipy 的dendrogram 函数中获取集群标签。

以下是几个选项/解决方法。

选项一

使用scipy的linkage和fcluster函数进行聚类并获取标签:

Z = sch.linkage(points, 'ward') # Note 'ward' is specified here to match the linkage used in sch.dendrogram.
labels = sch.fcluster(Z, t=10, criterion='distance') # t chosen to return two clusters.

# Cluster 1
np.where(labels == 1)

输出:(array([0, 3]),)

# Cluster 2
np.where(labels == 2)

输出:(array([1, 2, 4]),)

选项二

修改您当前对sklearn 的使用以返回两个集群:

hc = AgglomerativeClustering(n_clusters=2, affinity='euclidean',linkage='ward') # Again, 'ward' is specified here to match the linkage in sch.dendrogram.
y_hc = hc.fit_predict(points)

# Cluster 1
np.where(y_hc == 0)

输出:(array([0, 3]),)

# Cluster 2
np.where(y_hc == 1)

输出:(array([1, 2, 4]),)

【讨论】:

    猜你喜欢
    • 2021-07-21
    • 1970-01-01
    • 2021-12-14
    • 2011-12-22
    • 2017-12-03
    • 2013-06-09
    • 2021-04-04
    • 2023-03-12
    • 2017-10-24
    相关资源
    最近更新 更多