【发布时间】:2020-08-28 09:35:43
【问题描述】:
我们正在尝试删除异常值,但却出现了无限循环
对于一个学校项目,我们(我和一个朋友)认为创建一个基于数据科学的工具是个好主意。为此,我们开始清理数据库(我不会在这里导入它,因为它太大了(xlsx file,csv file))。我们现在尝试使用“标准差*3 + 平均值”规则删除“duration_minutes”列的异常值。
这是我们用来计算标准差和平均值的代码:
def calculateSD(database, column):
column = database[[column]]
SD = column.std(axis=None, skipna=None, level=None, ddof=1, numeric_only=None)
return SD
def calculateMean(database, column):
column = database[[column]]
mean = column.mean()
return mean
我们打算做以下事情:
#Now we have to remove the outliers using the code from the SD.py and SDfunction.py files
minutes = trainsData['duration_minutes'].tolist() #takes the column duration_minutes and puts it in a list
SD = int(calculateSD(trainsData, 'duration_minutes')) #calculates the SD of the column
mean = int(calculateMean(trainsData, 'duration_minutes'))
SDhigh = mean+3*SD
上面的代码计算起始值。然后我们开始一个while循环来删除异常值。删除异常值后,我们再次重新计算标准差、均值和 SDhigh。这是while循环:
while np.any(i >= SDhigh for i in minutes): #used to be >=, it doesnt matter for the outcome
trainsData = trainsData[trainsData['duration_minutes'] < SDhigh] #used to be >=, this caused an infinite loop so I changed it to <=. Then to <
minutes = trainsData['duration_minutes'].tolist()
SD = int(calculateSD(trainsData, 'duration_minutes')) #calculates the SD of the column
mean = int(calculateMean(trainsData, 'duration_minutes'))
SDhigh = mean+3*SD
print(SDhigh) #to see how the values changed and to confirm it is an infinite loop
输出如下:
611
652
428
354
322
308
300
296
296
296
296
它继续打印 296,经过数小时的尝试解决它,我们得出的结论是,我们并不像我们希望的那样聪明。
TL;DR:我们正在尝试删除所有高于标准差*3+mean 的值,直到没有留下任何值(我们每次都重新计算以检查是否还有异常值)。但是,我们得到了一个无限循环。
【问题讨论】:
-
一个潜在的问题可能是您将 mean 和 std 转换为 int ,这将改变值。尝试改用浮点数。
-
尝试检查删除前后的最大值是多少(并标记您的输出,以便您知道是什么)。
-
谢谢!虽然将类型改回浮点数给了我一些错误,但显然不需要整个 while 循环——我们试图使用的规则只需要应用一次。我们非常感谢您的快速遮阳篷,非常感谢
标签: python pandas csv infinite-loop standard-deviation