【问题标题】:Finding the most correlated item找到最相关的项目
【发布时间】:2019-07-28 03:50:06
【问题描述】:

我的餐厅销售详情如下。

+----------+------------+---------+----------+
| Location | Units Sold | Revenue | Footfall |
+----------+------------+---------+----------+
| Loc - 01 |        100 | 1,150   |       85 |
+----------+------------+---------+----------+

我想从下表餐厅数据中找到与上述最相关的餐厅

+----------+------------+---------+----------+
| Location | Units Sold | Revenue | Footfall |
+----------+------------+---------+----------+
| Loc - 02 |        100 | 1,250   |       60 |
| Loc - 03 |         90 | 990     |       90 |
| Loc - 04 |        120 | 1,200   |       98 |
| Loc - 05 |        115 | 1,035   |       87 |
| Loc - 06 |         89 | 1,157   |       74 |
| Loc - 07 |        110 | 1,265   |       80 |
+----------+------------+---------+----------+

请指导我如何使用 python 或 pandas 完成此操作。 注意:- 相关性表示在Units Sold、Revenue 和Footfall 方面最匹配/相似的餐厅。

【问题讨论】:

  • 你试过什么?请展示你的努力。
  • 我是熊猫新手。这就是为什么我要求指导我完成解决这个问题的过程。
  • 至少将您的数据读取为熊猫数据框,并计算相关矩阵。这应该是一个好的开始。
  • 根据哪个特征关联?
  • 你熟悉 Numpy 吗?我最近处理了一个类似的问题,我可以用 numpy/pandas 的方法为你解决这个问题。

标签: python pandas


【解决方案1】:

如果您的相关性应该被描述为最小欧几里德距离,则解决方案是:

#convert columns to numeric
df1['Revenue'] = df1['Revenue'].str.replace(',','').astype(int)
df2['Revenue'] = df2['Revenue'].str.replace(',','').astype(int)

#distance of all columns subtracted by first row of first DataFrame
dist = np.sqrt((df2['Units Sold']-df1.loc[0, 'Units Sold'])**2 + 
               (df2['Revenue']- df1.loc[0, 'Revenue'])**2 + 
               (df2['Footfall']- df1.loc[0, 'Footfall'])**2)

print (dist)
0    103.077641
1    160.390149
2     55.398556
3    115.991379
4     17.058722
5    115.542200
dtype: float64

#get index of minimal value and select row of second df
print (df2.loc[[dist.idxmin()]])
   Location  Units Sold  Revenue  Footfall
4  Loc - 06          89     1157        74

【讨论】:

  • @Datanovice - 谢谢,我希望这是 OP 需要的 ;)
  • @Datanovice @piRSquared 我得到了答案。谢谢你们。但是,我有一个后续问题。在这里,我想根据Loc - 01 选择所有相关的餐厅。如果我想选择所有类似相关的餐厅,而没有像Loc - 01 这样的基本餐厅怎么办?一种基于彼此相关性的聚类?
  • @Tommy - 我认为最好的办法是创建新问题。
【解决方案2】:

这可能是一个更好的方法,但我认为这是可行的,它非常冗长,所以我试图保持代码的清洁和可读性:

首先,让我们使用来自this 帖子的自定义 numpy 函数。

import numpy as np
import pandas as pd


def find_nearest(array, value):
    array = np.asarray(array)
    idx = (np.abs(array - value)).argmin()
    return array[idx]

然后使用您的数据框的数组,传入您的第一个数据框的值以找到最接近的匹配项。

us = find_nearest(df2['Units Sold'],df['Units Sold'][0])
ff = find_nearest(df2['Footfall'],df['Footfall'][0])
rev = find_nearest(df2['Revenue'],df['Revenue'][0])

print(us,ff,rev,sep=',')
100,87,1157

然后返回一个包含所有三个条件的数据框

    new_ df = (df2.loc[
    (df2['Units Sold'] == us) |
    (df2['Footfall'] == ff) |
    (df2['Revenue'] == rev)])

这给了我们:

    Location    Units Sold  Revenue Footfall
0   Loc - 02    100         1250    60
3   Loc - 05    115         1035    87
4   Loc - 06    89          1157    74

【讨论】:

    【解决方案3】:

    修复数据

    对于数字列。我可能过于笼统地概括了这一点。另外,我将索引设置为'Location' 列

    def fix(d):
        d.update(
            d.astype(str).replace(',', '', regex=True)
             .apply(pd.to_numeric, errors='ignore')
        )
        d.set_index('Location', inplace=True)
    
    fix(df1)
    fix(df2)
    

    曼哈顿距离

    df2.loc[[df2.sub(df1.loc['Loc - 01']).abs().sum(1).idxmin()]]
    
              Units Sold Revenue  Footfall
    Location                              
    Loc - 06          89    1157        74
    

    欧几里得距离

    df2.loc[[df2.sub(df1.loc['Loc - 01']).pow(2).sum(1).pow(.5).idxmin()]]
    
              Units Sold Revenue  Footfall
    Location                              
    Loc - 06          89    1157        74
    

    【讨论】:

      猜你喜欢
      • 2019-12-05
      • 1970-01-01
      • 2014-01-02
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多