【问题标题】:r remove outliers from a list of data.frames and make a new list of data.frames?r 从 data.frames 列表中删除异常值并创建一个新的 data.frames 列表?
【发布时间】:2017-05-08 15:37:06
【问题描述】:

我在 data.frame 中有一个 6 个列表

它有 3 列:

id、T_C、销售

T_C 是 TEST 或 CONTROL

有人在这里帮助了我,我学会了如何通过循环而不是单独的语句来找到 mean() 和 sd()。

现在我的目标是从 6 个列表中删除异常值并生成 6 个列表(在删除异常值之后)。

str(dfList) # 这是data.frames中的6个列表

我可以像这样得到每个列表的 mean() 和 sd():

list_mean_sd <- lapply(dfList,
                       function(df) 
                        {
                         df %>%
                           group_by(TC_INDICATOR) %>%
                           summarise(mean = mean(NET_SPEND),
                                     sd = sd(NET_SPEND))
                        })

> str(list_mean_sd)
List of 6  (1 obs. of  2 variables:)

我可以单独选择它们作为平均值或标准差:

sapply(list_mean_sd, "[", "mean")
sapply(list_mean_sd, "[", "sd")

基本上,我的目标是识别异常值并删除它们,生成替代集或后集。

**outliers are:  mean - 3*sd()  or  mean + 3*sd()

我已经完成了,但需要更多手动步骤,希望学习如何循环遍历这些集合和类似的东西,提前感谢您的帮助!

【问题讨论】:

    标签: r function functional-programming outliers identify


    【解决方案1】:

    试一试。首先,我创建数据,将其拆分为六个数据框,这些数据框位于一个列表中。

    set.seed(0)
    test_data <- data.frame(id = 1:10000, 
                            T_C = sample(c(TRUE, FALSE), size = 10000, replace = TRUE),
                            Sales = rnorm(n = 10000),
                            grp = sample(c("a", "b", "c", "d", "e", "f"), 
                                         size = 10000, replace = TRUE))
    
    test_split <- split(test_data, test_data$grp)
    

    然后,我在此列表中使用lapply 来识别我所称的z_scores,其计算方式为Sales 的mean 与每个Sales 之间的差异除以@987654327 @Sales。最后,我们对它们使用过滤器,以提取出z_score 绝对值大于 3 的那些。

    library(dplyr)
    outlier_list <- lapply(test_split, 
           function(m) group_by(m, T_C) %>% mutate(z_score = (Sales - mean(Sales)) / sd(Sales)) %>%
             ungroup() %>% filter(abs(z_score) >= 3)
    )
    
    > outlier_list
    $a
    # A tibble: 5 × 5
         id   T_C     Sales    grp   z_score
      <int> <lgl>     <dbl> <fctr>     <dbl>
    1   468  TRUE -2.995332      a -3.073314
    2  3026  TRUE  3.028495      a  3.075258
    3  5188  TRUE -3.097847      a -3.177952
    4  7993 FALSE -3.571076      a -3.823983
    5  9105  TRUE -3.216710      a -3.299276
    
    $b
    # A tibble: 6 × 5
         id   T_C     Sales    grp   z_score
      <int> <lgl>     <dbl> <fctr>     <dbl>
    1   264  TRUE  3.003494      b  3.003329
    2  2172  TRUE  3.001475      b  3.001326
    3  2980 FALSE -3.176356      b -3.222782
    4  3366 FALSE  3.009292      b  3.048559
    5  7477 FALSE  3.348301      b  3.392265
    6  7583  TRUE -3.089758      b -3.040911
    
    $c
    # A tibble: 2 × 5
         id   T_C    Sales    grp  z_score
      <int> <lgl>    <dbl> <fctr>    <dbl>
    1  8078  TRUE 3.015343      c 3.129923
    2  8991 FALSE 3.113526      c 3.058302
    
    $d
    # A tibble: 5 × 5
         id   T_C     Sales    grp   z_score
      <int> <lgl>     <dbl> <fctr>     <dbl>
    1   544  TRUE  3.289070      d  3.168235
    2  3791 FALSE  3.791938      d  3.769810
    3  6771 FALSE -3.157741      d -3.166861
    4  7864  TRUE  3.164128      d  3.045728
    5  9371  TRUE -3.026884      d -3.024655
    
    $e
    # A tibble: 6 × 5
         id   T_C     Sales    grp   z_score
      <int> <lgl>     <dbl> <fctr>     <dbl>
    1   186 FALSE  3.021541      e  3.046079
    2  1211  TRUE  3.414337      e  3.343521
    3  1665  TRUE  3.546282      e  3.473614
    4  3765 FALSE  3.363641      e  3.391142
    5  4172  TRUE  3.348820      e  3.278923
    6  7973 FALSE -2.987790      e -3.015284
    
    $f
    # A tibble: 6 × 5
         id   T_C     Sales    grp   z_score
      <int> <lgl>     <dbl> <fctr>     <dbl>
    1  1089  TRUE -3.195090      f -3.189979
    2  2452 FALSE  3.287591      f  3.212317
    3  3486 FALSE -3.334942      f -3.367962
    4  4198 FALSE -3.102578      f -3.137082
    5  8183  TRUE  3.081077      f  3.075324
    6  8656  TRUE  3.253873      f  3.247822
    

    显然,这只会给你异常值。如果您只想保留内点,请将&gt;= 3 更改为&lt; 3。

    更新为对内点进行 Wilcox 测试

    inlier_list <- lapply(test_split, 
                           function(m) group_by(m, T_C) %>% 
                            mutate(z_score = (Sales - mean(Sales)) / sd(Sales)) %>%
                             ungroup() %>% filter(abs(z_score) < 3)
    )
    

    我们只是使用 OP 评论中提到的参数在内点列表上运行 lapply。

    wilcox_test_res <- lapply(inlier_list, 
                              function(m) wilcox.test(m$Sales ~ m$T_C, 
                                                      mu= mean(m$Sales[m$T_C == TRUE]), 
                                                      conf.level=0.95,
    

    【讨论】:

    • 谢谢!我要到周二才上班去试一试。集团是干什么用的?我上了一门关于 R 的课程,但我发现我真的必须用它来获得任何好处,这门课在太短的时间内涵盖了太多内容。
    • 在这种情况下,group_by 允许我通过每个 data.frame 中的 T_C 列聚合 mean 和 sd。由于我在group_by 之后使用mutate,因此我使用mutate 添加的列是基于该分组计算的。
    • 不,我的意思是,您添加了一列 grp // 那是什么?
    • 这只是我制作的一个虚拟变量,因此我可以将模拟数据分成六组。我使用“grp”列来做到这一点。
    • 这很好用,我如何在 Wilcox.test 中使用循环,以便一次查看所有这 6 个列表?
    猜你喜欢
    • 2020-06-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-02-23
    • 2020-02-04
    • 1970-01-01
    • 2022-01-11
    • 2020-02-13
    相关资源
    最近更新 更多