【问题标题】:Speeding up outliers check on a pandas Series加快对熊猫系列的异常值检查
【发布时间】:2013-02-10 19:00:31
【问题描述】:

我正在使用不同的标准偏差标准对 pandas Series 对象进行两次通过的异常值检查。但是,我为此使用了两个循环,并且运行速度非常慢。不知道有没有什么pandas“技巧”可以加快这一步。

这是我正在使用的代码(警告非常丑陋的代码!):

def find_outlier(point, window, n):
    return np.abs(point - nanmean(window)) >= n * nanstd(window)

def despike(self, std1=2, std2=20, block=100, keep=0):
    res = self.values.copy()
    # First run with std1:
    for k, point in enumerate(res):
        if k <= block:
            window = res[k:k + block]
        elif k >= len(res) - block:
            window = res[k - block:k]
        else:
            window = res[k - block:k + block]
        window = window[~np.isnan(window)]
        if np.abs(point - window.mean()) >= std1 * window.std():
            res[k] = np.NaN
    # Second run with std2:
    for k, point in enumerate(res):
        if k <= block:
            window = res[k:k + block]
        elif k >= len(res) - block:
            window = res[k - block:k]
        else:
            window = res[k - block:k + block]
        window = window[~np.isnan(window)]
        if np.abs(point - window.mean()) >= std2 * window.std():
            res[k] = np.NaN
    return Series(res, index=self.index, name=self.name)

【问题讨论】:

    标签: python pandas outliers


    【解决方案1】:

    我不确定你在用那个块做什么,但是在系列中查找异常值应该很容易:

    In [1]: s > s.std() * 3
    

    其中 s 是您的系列,3 是离群值状态要超过多少标准差。此表达式将返回一系列布尔值,然后您可以通过以下方式索引该系列:

    In [2]: s.head(10)
    Out[2]:
    0    1.181462
    1   -0.112049
    2    0.864603
    3   -0.220569
    4    1.985747
    5    4.000000
    6   -0.632631
    7   -0.397940
    8    0.881585
    9    0.484691
    Name: val
    
    In [3]: s[s > s.std() * 3]
    Out[3]:
    5    4
    Name: val
    

    更新:

    解决关于块的评论。我认为在这种情况下您可以使用pd.rolling_std():

    In [53]: pd.rolling_std(s, window=5).head(10)
    Out[53]:
    0         NaN
    1         NaN
    2         NaN
    3         NaN
    4    0.871541
    5    0.925348
    6    0.920313
    7    0.370928
    8    0.467932
    9    0.391485
    
    In [55]: abs(s) > pd.rolling_std(s, window=5) * 3
    
    Docstring:
    Unbiased moving standard deviation
    
    Parameters
    ----------
    arg : Series, DataFrame
    window : Number of observations used for calculating statistic
    min_periods : int
        Minimum number of observations in window required to have a value
    freq : None or string alias / date offset object, default=None
        Frequency to conform to before computing statistic
        time_rule is a legacy alias for freq
    
    Returns
    -------
    y : type of input argument
    

    【讨论】:

    • 嗨 Zelazny7。该块是因为我需要将每个点与远离它的 100 点进行比较,而不是整个系列。这就是我需要循环的原因。
    • 谢谢,这正是我所需要的。
    • 请注意,此解决方案假设数据以零为中心。一个更准确的答案:abs(s - s.mean()) > pd.rolling_std(s, window=5) * 3
    猜你喜欢
    • 1970-01-01
    • 2019-01-14
    • 1970-01-01
    • 2020-10-05
    • 1970-01-01
    • 2022-10-16
    • 2017-07-30
    相关资源
    最近更新 更多