【问题标题】:FInding K-mean distance寻找K-mean距离
【发布时间】:2020-06-14 05:53:29
【问题描述】:

我有一个包含 13 个特征和 1000 万行的数据库。我想应用 k-mean 来消除任何异常。我的想法是应用k-mean,创建一个包含数据点和聚类质心之间距离的新列,以及一个包含平均距离的新列,如果距离大于平均距离,我删除整行。但是我写的代码好像不行。

数据集示例: https://drive.google.com/open?id=1iB1qjnWQyvoKuN_Pa8Xk4BySzXVTwtUk

df = pd.read_csv('Final After Simple Filtering.csv',index_col=None,low_memory=True)


# Dropping columns with low feature importance
del df['AmbTemp_DegC']
del df['NacelleOrientation_Deg']
del df['MeasuredYawError']



#applying kmeans
#applying kmeans
kmeans = KMeans( n_clusters=8)


clusters= kmeans.fit_predict(df)

centroids = kmeans.cluster_centers_

distance1 = kmeans.fit_transform(df)

distance2 = distance1.mean()

df['distances']=distance1-distance2

df = df[df['distances'] >=0]


del df['distances']

df.to_csv('/content//drive/My Drive/K TEST.csv', index=False)

错误:

KeyError                                  Traceback (most recent call last)
/usr/local/lib/python3.6/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance)
   2896             try:
-> 2897                 return self._engine.get_loc(key)
   2898             except KeyError:

pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'distances'

During handling of the above exception, another exception occurred:

KeyError                                  Traceback (most recent call last)
9 frames
pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/index.pyx in pandas._libs.index.IndexEngine.get_loc()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'distances'

During handling of the above exception, another exception occurred:

ValueError                                Traceback (most recent call last)
/usr/local/lib/python3.6/dist-packages/pandas/core/internals/blocks.py in __init__(self, values, placement, ndim)
    126             raise ValueError(
    127                 "Wrong number of items passed {val}, placement implies "
--> 128                 "{mgr}".format(val=len(self.values), mgr=len(self.mgr_locs))
    129             )
    130 

ValueError: Wrong number of items passed 8, placement implies 1

谢谢

【问题讨论】:

  • 您能告诉我们您遇到了什么错误吗?
  • @Ehrendil 我已在主帖中发布了错误。
  • 我建议也发布一个数据框样本
  • 我已将我的数据集样本添加到主帖

标签: python pandas dataframe machine-learning jupyter-notebook


【解决方案1】:

这是您上一个问题的后续回答。

import seaborn as sns
import pandas as pd
titanic = sns.load_dataset('titanic')
titanic = titanic.copy()
titanic = titanic.dropna()
titanic['age'].plot.hist(
  bins = 50,
  title = "Histogram of the age variable"
)


from scipy.stats import zscore
titanic["age_zscore"] = zscore(titanic["age"])
titanic["is_outlier"] = titanic["age_zscore"].apply(
  lambda x: x <= -2.5 or x >= 2.5
)
titanic[titanic["is_outlier"]]


ageAndFare = titanic[["age", "fare"]]
ageAndFare.plot.scatter(x = "age", y = "fare")


from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
ageAndFare = scaler.fit_transform(ageAndFare)
ageAndFare = pd.DataFrame(ageAndFare, columns = ["age", "fare"])
ageAndFare.plot.scatter(x = "age", y = "fare")


from sklearn.cluster import DBSCAN
outlier_detection = DBSCAN(
  eps = 0.5,
  metric="euclidean",
  min_samples = 3,
  n_jobs = -1)
clusters = outlier_detection.fit_predict(ageAndFare)

clusters


from matplotlib import cm
cmap = cm.get_cmap('Accent')
ageAndFare.plot.scatter(
  x = "age",
  y = "fare",
  c = clusters,
  cmap = cmap,
  colorbar = False
)

查看此链接了解所有详情。

https://www.mikulskibartosz.name/outlier-detection-with-scikit-learn/

在今天之前我从未听说过“局部异常因子”。当我用谷歌搜索它时,我得到了一些似乎表明它是 DBSCAN 的衍生产品的信息。最后,我认为我的第一个答案实际上是检测异常值的最佳方法。 DBSCAN 是一种聚类算法,它碰巧发现异常值,这些异常值实际上被认为是“噪声”。我不认为 DBSCAN 的主要目的是异常检测,而是聚类。总之,正确选择超参数需要一些技巧。此外,DBSCAN 在非常大的数据集上可能会很慢,因为它隐含地需要计算每个样本点的经验密度,导致二次最坏情况时间复杂度,这在大型数据集上非常慢。

【讨论】:

    【解决方案2】:

    你:我想应用 k-mean 来消除任何异常。

    实际上,KMeas 会检测异常​​并将它们包含在最近的集群中。损失函数是从每个点到其分配的簇质心的最小平方和。如果您想剔除异常值,请考虑使用 z-score 方法。

    import numpy as np
    import pandas as pd
    
    # import your data
    df = pd.read_csv('C:\\Users\\your_file.csv)
    
    # get only numerics
    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']
    newdf = df.select_dtypes(include=numerics)
    
    df = newdf
    
    # count rows in DF before kicking out records with z-score over 3
    df.shape
    
    # handle NANs
    df = df.fillna(0)
    
    
    from scipy import stats
    df = df[(np.abs(stats.zscore(df)) < 3).all(axis=1)]
    df.shape
    
    
    df = pd.DataFrame(np.random.randn(100, 3))
    from scipy import stats
    df[(np.abs(stats.zscore(df)) < 3).all(axis=1)]
    
    # count rows in DF before kicking out records with z-score over 3
    df.shape
    

    另外,有空的时候看看这些链接。

    https://medium.com/analytics-vidhya/effect-of-outliers-on-k-means-algorithm-using-python-7ba85821ea23

    https://statisticsbyjim.com/basics/outliers/

    【讨论】:

    • 我已经在 K-mean 上阅读了几天,我阅读的大部分内容并不是用于异常检测,并且有更好的方法可以做到这一点。不幸的是,我最后一年项目的一部分是比较孤立森林、椭圆包络和 K-mean 来去除异常值。也非常感谢您的帮助。
    • 也看看 DBSCAN。该算法将帮助您找到实际的异常值,我认为这些异常值被定义为“噪声”。 hdbscan.readthedocs.io/en/latest/outlier_detection.html & hdbscan.readthedocs.io/en/latest/parameter_selection.html
    • 如果我发现很难使用 K-mean 从我的数据集中删除异常值,使用 DBSCAN 或局部异常值因子 (LOF) 会更好吗?
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2016-03-18
    • 2017-08-03
    • 2019-04-27
    • 2018-12-17
    • 2018-01-15
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多