【问题标题】:An infinite while loop in python with pandas calculating the standard deviationpython中的无限while循环,熊猫计算标准偏差
【发布时间】:2020-08-28 09:35:43
【问题描述】:

我们正在尝试删除异常值,但却出现了无限循环

对于一个学校项目,我们(我和一个朋友)认为创建一个基于数据科学的工具是个好主意。为此,我们开始清理数据库(我不会在这里导入它,因为它太大了(xlsx file,csv file))。我们现在尝试使用“标准差*3 + 平均值”规则删除“duration_minutes”列的异常值。

这是我们用来计算标准差和平均值的代码:

def calculateSD(database, column):
    column = database[[column]]
    SD = column.std(axis=None, skipna=None, level=None, ddof=1, numeric_only=None)
    return SD

def calculateMean(database, column):
    column = database[[column]]
    mean = column.mean()
    return mean

我们打算做以下事情:

#Now we have to remove the outliers using the code from the SD.py and SDfunction.py files
minutes = trainsData['duration_minutes'].tolist() #takes the column duration_minutes and puts it in a list
SD = int(calculateSD(trainsData, 'duration_minutes')) #calculates the SD of the column
mean = int(calculateMean(trainsData, 'duration_minutes'))
SDhigh = mean+3*SD

上面的代码计算起始值。然后我们开始一个while循环来删除异常值。删除异常值后,我们再次重新计算标准差、均值和 SDhigh。这是while循环:

while np.any(i >= SDhigh for i in minutes): #used to be >=, it doesnt matter for the outcome
    trainsData = trainsData[trainsData['duration_minutes'] < SDhigh] #used to be >=, this caused an infinite loop so I changed it to <=. Then to <
    minutes = trainsData['duration_minutes'].tolist()
    SD = int(calculateSD(trainsData, 'duration_minutes')) #calculates the SD of the column
    mean = int(calculateMean(trainsData, 'duration_minutes'))
    SDhigh = mean+3*SD
    print(SDhigh) #to see how the values changed and to confirm it is an infinite loop

输出如下:

611
652
428
354
322
308
300
296
296
296
296

它继续打印 296,经过数小时的尝试解决它,我们得出的结论是,我们并不像我们希望的那样聪明。


TL;DR:我们正在尝试删除所有高于标准差*3+mean 的值,直到没有留下任何值(我们每次都重新计算以检查是否还有异常值)。但是,我们得到了一个无限循环。

【问题讨论】:

  • 一个潜在的问题可能是您将 mean 和 std 转换为 int ,这将改变值。尝试改用浮点数。
  • 尝试检查删除前后的最大值是多少(并标记您的输出,以便您知道是什么)。
  • 谢谢!虽然将类型改回浮点数给了我一些错误,但显然不需要整个 while 循环——我们试图使用的规则只需要应用一次。我们非常感谢您的快速遮阳篷,非常感谢

标签: python pandas csv infinite-loop standard-deviation


【解决方案1】:

你让事情变得比他们必须的更困难。计算标准偏差以去除异常值然后重新计算它等过于复杂(并且在统计上不合理)。你最好使用百分位数而不是标准差

import numpy as np
import pandas as pd

# create data
nums = np.random.normal(50, 8, 200)
df = pd.DataFrame(nums, columns=['duration'])

# set threshold based on percentiles
threshold = df['duration'].quantile(.95) * 2

# now only keep rows that are below the threshold
df = df[df['duration']<threshold]

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2018-11-17
    • 2023-03-29
    • 1970-01-01
    • 1970-01-01
    • 2021-06-30
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多