【问题标题】:Identifying chunks of outliers from a 1D and 2D data in Python在 Python 中从一维和二维数据中识别异常值块
【发布时间】:2015-09-28 11:12:54
【问题描述】:

数据:我在一列中有一个数据 d,该数据随其他两个变量的变化而变化, a 和 b,在其他两列中定义。我的目标是识别 d 中的块或异常值。这些异常值块似乎不是异常值,但就我而言,我想识别那些不落在可以拟合线性线的数据云上的数据。

问题:尽管我以前从未做过聚类分析,但这个名字听起来好像可以实现我想要做的事情。如果我选择进行聚类分析,我想针对以下两种情况进行:

  1. 与 a 和 d
  2. 与 a、b 和 d >

我做了一些搜索,发现#1,使用KernelDensity 模块会更合适,而对于#2,使用MeahShift 模块将是一个不错的选择,无论是在 Python 中。

问题:我以前从未做过聚类分析,所以我无法理解他们文档中给出的KernelDensity 和MeahShift 的示例(分别为here 和here) .有人可以解释一下我如何使用 KernelDensity 和 MeahShift 来识别 d 中的异常值“块”以用于案例 1 和 2?

【问题讨论】:

  • 我认为你首先需要一个稳健的回归,因为你的数据已经被一些异常值污染了。一旦拟合了稳健的回归,则在每个点计算的均方误差可以用作到聚类中心(回归线)的距离度量。那些具有大 MSE 的观察结果可能是异常值。
  • sklearn 中稳健回归的参考链接。 scikit-learn.org/stable/modules/…
  • @JianxunLi:对不起,我看不懂那个参考文献中给出的例子。你能举个简单的例子吗?
  • 代码似乎适用于您的示例数据文件。我已经在我的帖子中加入了,请看一下。

标签: python scikit-learn cluster-analysis outliers chunks


【解决方案1】:

首先,KernelDensity 用于非参数方法。由于您坚信关系是线性的(即参数模型),KernelDensity 不是此任务中最合适的选择。

以下是识别异常值的示例代码。

import matplotlib.pyplot as plt
import numpy as np
from sklearn.linear_model import RANSACRegressor


# data: 1000 obs, 100 of them are outliers
# =====================================================
np.random.seed(0)
a = np.random.randn(1000)
b = np.random.randn(1000)
d = 2 * a - b + np.random.randn(1000)
# the last 100 are outliers
d[-100:] = d[-100:] + 10 * np.abs(np.random.randn(100))

fig, axes = plt.subplots(ncols=2, sharey=True)
axes[0].scatter(a, d, c='g')
axes[0].set_xlabel('a')
axes[0].set_ylabel('d')
axes[1].scatter(b, d, c='g')
axes[1].set_xlabel('b')

# processing
# =====================================================
# robust regression
robust_estimator = RANSACRegressor(random_state=0)
robust_estimator.fit(np.vstack([a,b]).T, d)
d_pred = robust_estimator.predict(np.vstack([a,b]).T)

# calculate mse
mse = (d - d_pred.ravel()) ** 2

# get 50 largest mse, 50 is just an arbitrary choice and it doesn't assume that we already know there are 100 outliers
index = argsort(mse)
fig, axes = plt.subplots(ncols=2, sharey=True)
axes[0].scatter(a[index[:-50]], d[index[:-50]], c='b', label='inliers')
axes[0].scatter(a[index[-50:]], d[index[-50:]], c='r', label='outliers')
axes[0].set_xlabel('a')
axes[0].set_ylabel('d')
axes[0].legend(loc='best')
axes[1].scatter(b[index[:-50]], d[index[:-50]], c='b', label='inliers')
axes[1].scatter(b[index[-50:]], d[index[-50:]], c='r', label='outliers')
axes[1].legend(loc='best')
axes[1].set_xlabel('b')

用于您的示例数据

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
from sklearn.linear_model import RANSACRegressor

df = pd.read_excel('/home/Jian/Downloads/Data.xlsx').dropna()

a = df.a.values.reshape(len(df), 1)
d = df.d.values.reshape(len(df), 1)

fig, axes = plt.subplots(ncols=2, sharey=True)
axes[0].scatter(a, d, c='g')
axes[0].set_xlabel('a')
axes[0].set_ylabel('d')

robust_estimator = RANSACRegressor(random_state=0)
robust_estimator.fit(a, d)
d_pred = robust_estimator.predict(a)

# calculate mse
mse = (d - d_pred) ** 2

index = np.argsort(mse.ravel())

axes[1].scatter(a[index[:-50]], d[index[:-50]], c='b', label='inliers', alpha=0.2)
axes[1].scatter(a[index[-50:]], d[index[-50:]], c='r', label='outliers')
axes[1].set_xlabel('a')
axes[1].legend(loc=2)

【讨论】:

  • 根据我在您的代码中的理解,您是-1)在您已经知道异常值的前提下开始编写代码,2)对包括异常值在内的所有数据进行回归拟合。另外,您在稳健回归中使用的参数T 是什么?但是,我的目标是 1)首先破译这些异常值块,2)仅对位于异常值下方的数据云进行线性回归。
  • @Pupil 不,我不认为对异常值有任何了解。后半部分的所有代码都不假设它知道异常值是最后 100 个 obs。上面的代码演示了如何删除异常值。如果您愿意,可以使用剩余的内点重新拟合线性回归。 .T 只是转置运算符,请确保每一列都是一个特征。
  • 你为什么要求代码将 50 个最大的 mse 标记为异常值?基本上,我不知道有多少异常值,也不知道它们是否存在。我希望代码能够识别那些异常值并在没有异常值的剩余数据上拟合一条线。另外,现在你已经包含了a和b这两个参数;如果我只包含一个参数,比如a,这会起作用吗?
  • @Pupil 1) d_pred 基于内点和异常值,但由于我们使用稳健回归,回归线对异常值是稳健的。如果您只想为另一个标准线性回归重新拟合内点,它将产生非常相似的回归线。 2)据我所知,很难精确检测数据集中有多少异常值,如果它的mse 比mse 的中位数大5 倍,你可以使用它作为异常值的标准。 3)您可以更改常量并预测新的d_pred,以便np.percentile(d_pre - d, 0.05) 为您提供0.0。
  • @Pupil 您需要解决a0 使您的功能np.percentile(d_pre - d, 0.05) = 0 的原因。我认为使用scipy 寻根算法是可行的。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-12-10
  • 2021-01-17
  • 2018-01-18
  • 1970-01-01
  • 2018-07-21
  • 1970-01-01
相关资源
最近更新 更多