【问题标题】:Marking outliers on a Scatter Plot在散点图上标记异常值
【发布时间】:2020-02-01 20:36:52
【问题描述】:

我有一个如下所示的数据框:

 print(df.head(10))

 day         CO2
   1  549.500000
   2  663.541667
   3  830.416667
   4  799.695652
   5  813.850000
   6  769.583333
   7  681.941176
   8  653.333333
   9  845.666667
  10  436.086957

然后我使用以下函数和代码行从 CO2 列中获取 ouliers:

def estimate_gaussian(dataset):

    mu = np.mean(dataset)#moyenne cf mu
    sigma = np.std(dataset)#écart_type/standard deviation
    limit = sigma * 1.5

    min_threshold = mu - limit
    max_threshold = mu + limit

    return mu, sigma, min_threshold, max_threshold

mu, sigma, min_threshold, max_threshold = estimate_gaussian(df['CO2'].values)


condition1 = (dataset < min_threshold)
condition2 = (dataset > max_threshold)

outliers1 = np.extract(condition1, dataset)
outliers2 = np.extract(condition2, dataset)

outliers = np.concatenate((outliers1, outliers2), axis=0)

这给了我以下结果:

print(outliers)

[830.41666667 799.69565217 813.85       769.58333333 845.66666667]

现在我想在散点图上用红色标记那些异常值。

您可以在下面找到到目前为止我用来在散点图上用红色标记单个异常值的代码,但我找不到为异常值列表的每个元素(即 numpy.ndarray)执行此操作的方法:

y = df['CO2']

x = df['day']

col = np.where(x<0,'k',np.where(y<845.66666667,'b','r'))

plt.scatter(x, y, c=col, s=5, linewidth=3)
plt.show()

这是我得到的,但我想要所有 ouliers 的相同结果。你能帮帮我吗?

https://ibb.co/Ns9V7Zz

【问题讨论】:

    标签: python matplotlib plot scatter-plot outliers


    【解决方案1】:

    可能不是最有效的解决方案,但我觉得多次调用plt.scatter 更容易,每次传递一个 xy 对。由于我们从不调用新图形(例如使用plt.figure()),因此每个 xy 对都绘制在同一个图形上。

    然后,在每次迭代中,我们只需要检查 y 值是否为异常值。如果是,我们更改 plt.scatter 调用中的 color 关键字参数。

    试试这个:

    mu, sigma, min_threshold, max_threshold = estimate_gaussian(df['CO2'].values)
    
    xs = df['day']
    ys = df['CO2']
    
    for x, y in zip(xs, ys):
        color = 'blue'  # non-outlier color
        if not min_threshold <= y <= max_threshold:  # condition for being an outlier
            color = 'red'  # outlier color
        plt.scatter(x, y, color=color)
    plt.show()
    

    【讨论】:

      【解决方案2】:

      您可以创建一个附加列(布尔值),在其中定义该点是否为异常值(真)或不是(假),然后使用两个散点图:

      df["outlier"] = # your boolean np array goes in here
      plt.scatter[df.loc[df["outlier"], "day"], df.loc[df["outlier"], "CO2"], color="k"]
      plt.scatter[df.loc[~df["outlier"], "day"], df.loc[~df["outlier"], "CO2"], color="r"]
      

      【讨论】:

        【解决方案3】:

        我不确定您的 col 列表背后的想法是什么,但您可以将 col 替换为

        col = ['red' if yy in list(outliers) else 'blue' for yy in y] 
        

        【讨论】:

          【解决方案4】:

          有几种方法,一种是根据您的条件创建一系列颜色并将其传递给c 参数。

          df = pd.DataFrame({'CO2': {0: 549.5,
            1: 663.54166699999996,
            2: 830.41666699999996,
            3: 799.695652,
            4: 813.85000000000002,
            5: 769.58333300000004,
            6: 681.94117599999993,
            7: 653.33333300000004,
            8: 845.66666699999996,
            9: 436.08695700000004},
           'day': {0: 1, 1: 2, 2: 3, 3: 4, 4: 5, 5: 6, 6: 7, 7: 8, 8: 9, 9: 10}})
          
          In [11]: colors = ['r' if n<750 else 'b' for n in df['CO2']]
          
          In [12]: colors
          Out[12]: ['r', 'r', 'b', 'b', 'b', 'b', 'r', 'r', 'b', 'r']
          
          In [13]: plt.scatter(df['day'],df['CO2'],c=colors)
          

          或者使用np.where创建序列

          In [14]: colors = np.where(df['CO2'] < 750, 'r', 'b')
          

          【讨论】:

            【解决方案5】:

            这是一个快速的解决方案:

            我将重新创建您已经开始的内容。您只共享了数据框的头部,但无论如何,我只是插入了一些随机异常值。看起来您的“estimate_gaussian()”函数只能返回两个异常值?

            import pandas as pd
            import matplotlib.pyplot as plt
            
            df = pd.DataFrame([549.500000,
                            50.0000000,
                            830.416667,
                            799.695652,
                            1200.00000,
                            769.583333,
                            681.941176,
                            1300.00000,
                            845.666667,
                            436.086957], 
                            columns=['CO2'],
                            index=list(range(1,11)))
            
            def estimate_gaussian(dataset):
            
                mu = np.mean(dataset) # moyenne cf mu
                sigma = np.std(dataset) # écart_type/standard deviation
                limit = sigma * 1.5
            
                min_threshold = mu - limit
                max_threshold = mu + limit
            
                return mu, sigma, min_threshold, max_threshold
            
            mu, sigma, min_threshold, max_threshold = estimate_gaussian(df.values)
            
            condition1 = (df < min_threshold)
            condition2 = (df > max_threshold)
            
            outliers1 = np.extract(condition1, df)
            outliers2 = np.extract(condition2, df)
            
            outliers = np.concatenate((outliers1, outliers2), axis=0)
            

            然后我们将绘制:

            df_red = df[df.values==outliers]
            
            plt.scatter(df.index,df.values)
            plt.scatter(df_red.index,df_red.values,c='red')
            plt.show()
            

            如果您需要更细微的东西,请告诉我!

            【讨论】:

            • 仅供参考,通过将 df_red = df[df.CO2.isin([outliers1, outliers2])] 替换为 df_red = df[df.values==outliers] 解决了问题
            猜你喜欢
            • 2013-05-01
            • 1970-01-01
            • 2019-09-30
            • 1970-01-01
            • 1970-01-01
            • 1970-01-01
            • 2013-02-17
            • 2021-11-06
            • 2012-03-18
            相关资源
            最近更新 更多