【问题标题】:Which clustering distance-metric to find the most correlated groups of items哪个聚类距离度量可以找到最相关的项目组
【发布时间】:2019-12-05 21:37:19
【问题描述】:

我有如下餐厅销售数据,并希望找到彼此相关的餐厅。我正在寻找一种基于彼此相关性的聚类;其中“相关性”是指“销售量、收入和客流量组合最匹配/相似的餐厅”。 (注:这是corelatedItems的后续问题)

+----------+------------+---------+----------+
| Location | Units Sold | Revenue | Footfall |
+----------+------------+---------+----------+
| Loc - 01 |        100 | 1,150   |       85 |
| Loc - 02 |        100 | 1,250   |       60 |
| Loc - 03 |         90 | 990     |       90 |
| Loc - 04 |        120 | 1,200   |       98 |
| Loc - 05 |        115 | 1,035   |       87 |
| Loc - 06 |         89 | 1,157   |       74 |
| Loc - 07 |        110 | 1,265   |       80 |
+----------+------------+---------+----------+

【问题讨论】:

  • 你想要的输出是什么?
  • 您要的是 distance metric used in clustering,请阅读该 sklearn 文档。
  • 你已经有一个good answer to the correlation question,剩下的只是“我如何在 sklearn 中进行聚类?”,这在 sklearn 文档中有介绍。请尝试编写您自己的(sklearn+pandas)代码,然后告诉我们您遇到的问题。
  • sklearn 是 Python 的主要机器学习库之一,请查看并浏览它的文档(分类器、功能、管道等),听起来你会经常使用它。
  • @smci 谢谢。将检查 sklearn 并试一试。

标签: python pandas cluster-analysis distance


【解决方案1】:

首先,将dataframe的索引设置为Location列,方便索引

df1 = df1.set_index('Location')

接下来,生成所有餐厅组合进行比较:

import itertools
pairs = list(itertools.combinations(df1.index.values, 2))

接下来,定义一个比较函数。让我们使用上一篇文章中使用的那个

import numpy as np
def compare_function(row1, row2):
    return np.sqrt((row1['Units Sold']-row2['Units Sold'])**2 + 
           (row1['Revenue']- row2['Revenue'])**2 + 
           (row1['Footfall']- row2.loc[0, 'Footfall'])**2)

接下来,遍历所有pair,得到比较函数的结果:

results = [(row1, row2, compare_function(df1.loc[row1], df1.loc[row2]))
      for row1, row2 in pairs]

您现在拥有所有成对餐厅及其彼此之间距离的列表。

【讨论】:

  • sklearn.cluster 多年前就实现了集群,无需重新发明轮子,而且有很多优秀的教程。
猜你喜欢
  • 2015-11-25
  • 1970-01-01
  • 2018-09-16
  • 2012-04-15
  • 2011-07-21
  • 2020-03-05
  • 2013-05-05
  • 2011-12-09
  • 2012-07-09
相关资源
最近更新 更多