【问题标题】:How calculate list python into matrix similarity如何将列表python计算为矩阵相似度
【发布时间】:2016-04-01 07:39:08
【问题描述】:

我的数据有问题

我使用 python 从我的数据库中读取数据,假设分配给变量data

type(data) 是list,实际上是list of list

data = [(1, 'Shirt', 2),(1, 'Pants', 3),(2, 'Top', 2),(2, 'Shirt', 1),(2, 'T-Shirt', 4), (3, 'Shirt', 3),(3, 'T-Shirt', 2)]

data[0][0] is unique_id 和 data[0][1] is category_product 和 data[0][2] is count

我需要基于category_product 使用余弦相似度计算unique_id 1 和2 之间的相似度(我计划使用scipy)

  • unique_id不只是两个,可以超过2个

我想我需要将我的data 转换为矩阵:

unique_id | Shirt | Pants | Top | T-Shirt
1 | 2 | 3 | 0 | 0 
2 | 1 | 0 | 2 | 4
3 | 3 | 0 | 0 | 2

我想用余弦相似度计算这个矩阵,输出是:

1,2,0.121045506534
1,3,0.461538461538
2,3,0.665750285936
  • Sim(1,2) = 0.121045506534

我如何用 python 做到这一点?

谢谢

【问题讨论】:

  • 你的数据不是列表的列表,而是元组的列表 :) 有很大的不同。
  • 啊,你说得对,对不起我的错误:)

标签: python numpy scipy


【解决方案1】:
import pandas as pd
from scipy import spatial
from itertools import combinations

df = pd.DataFrame(data, columns=['unique_id', 'category_product', 'count'])

pt = df.pivot(index='unique_id', columns='category_product', values='count').fillna(0)

>>> pt
category_product  Pants  Shirt  T-Shirt  Top
unique_id                                   
1                     3      2        0    0
2                     0      1        4    2
3                     0      3        2    0

combos = combinations(pt.index, 2)
>>> [(a, b, 1 - spatial.distance.cosine(pt.ix[a].values, pt.ix[b].values)) 
     for a, b in combos]
[(1, 2, 0.12104550653376045),
 (1, 3, 0.46153846153846168),
 (2, 3, 0.66575028593568275)]

【讨论】:

  • 也许需要import pandas as pd 对吧?谢谢@alexander
  • 嗨@Alexander,对不起,如果我的真实数据分配给pt,它变成了[4260 rows x 248 columns],如何让它更快地计算?谢谢 :)
  • 拥有 4260 个唯一 ID 值,大约有 (4260^2) / 2 种可能的配对。加快速度的唯一方法是获得更快的计算机或减少唯一 ID 的数量。还有其他解决方案,但它们非常专业,超出了这个论坛。
猜你喜欢
  • 2016-10-22
  • 2017-11-07
  • 1970-01-01
  • 2018-08-01
  • 2019-04-22
  • 2017-06-17
  • 1970-01-01
  • 2016-02-15
  • 2014-03-25
相关资源
最近更新 更多