【问题标题】:missing values from grouped data R分组数据 R 中的缺失值
【发布时间】:2018-05-09 15:43:37
【问题描述】:

如果这是一个简单的问题,我们深表歉意。我有整齐(长)格式的数据。我想看看Factor Name 中的值集对于Sample Name 中的每个样本有何不同。我相信使用 group_by 函数是可能的。

# Groups:   Sample Name
  `Sample Name` `Factor Name`    mean
   <fct>         <fct>           <dbl>
 1 S1            ABCD            -5.15
 2 S1            EFGH             7.74
 3 S1            IJKL            -7.43
 4 S2            ABCD             4.35
 5 S2            EFGH            -2.15
 6 S2            IJKL             2.33
 7 S3            ABCD             5.53
 8 S3            EFGH             2.84
 9 S3            IJKL             1.61
10 S3            MNOP             NaN   

我也尝试过聚合,虽然它给出了输出,但我更喜欢 group_by 或管道有效的方法。

Aggregate(`Factor Name` ~ `Sample Name`, df, FUN= function(x) setdiff(unique(df$`Factor Name`),x))

如果可能的话,我希望能够为每个示例名称添加缺少的Factor Name,如下所示:

# Groups:   Sample Name
  `Sample Name` `Factor Name`    mean
   <fct>         <fct>           <dbl>
 1 S1            ABCD            -5.15
 2 S1            EFGH             7.74
 3 S1            IJKL            -7.43
 4 S1            MNOP             NaN
 5 S2            ABCD             4.35
 6 S2            EFGH            -2.15
 7 S2            IJKL             2.33
 8 S2            MNOP             NaN
 9 S3            ABCD             5.53
10 S3            EFGH             2.84
11 S3            IJKL             1.61
12 S3            MNOP             NaN   

【问题讨论】:

标签: r dplyr


【解决方案1】:

tidyr::expand 和 tidyr::compelete 函数对您想要实现的目标非常有用。

加载包:

library(dplyr)
library(tidyr)

创建一个虚拟数据集:

df <- data_frame(sample_name = factor(c(rep(c('S1', 'S2', 'S3'), each = 3), 'S3')),
                 factor_name = factor(c(rep(c('ABCD', 'EFGH', 'IJKL'), 3), 'MNOP')),
                 mean = rnorm(n = 10, sd = 10))

问题 1

为sample_name 中的每个样本获取factor_name 中的一组值的差异:

# Return ONLY those levels of sample_name that are missing a level of factor_name
df %>% 
    # Expand to all unique combinations
    expand(sample_name, factor_name) %>% 
    # Extract the difference
    setdiff(., select(df, -mean)) 

#> # A tibble: 2 x 2
#>   sample_name factor_name
#>   <fct>       <fct>      
#> 1 S1          MNOP       
#> 2 S2          MNOP

# Return ALL levels of sample_name, along with any missing levels of factor_name
df %>% 
    # Expand to all unique combinations
    expand(sample_name, factor_name) %>% 
    # Extract the difference
    setdiff(., select(df, -mean)) %>% 
    # Expand to show all levels of sample_name
    complete(sample_name)

#> # A tibble: 3 x 2
#>   sample_name factor_name
#>   <fct>       <fct>      
#> 1 S1          MNOP       
#> 2 S2          MNOP       
#> 3 S3          <NA>

问题 2

为每个sample_name 添加缺少的factor_name:

# Expand to include ALL levels of factor_name within sample_name
df %>% 
    complete(sample_name, factor_name) 

#> # A tibble: 12 x 3
#>    sample_name factor_name     mean
#>    <fct>       <fct>          <dbl>
#>  1 S1          ABCD         16.6   
#>  2 S1          EFGH         -0.0803
#>  3 S1          IJKL          4.80  
#>  4 S1          MNOP         NA     
#>  5 S2          ABCD          3.80  
#>  6 S2          EFGH         -1.24  
#>  7 S2          IJKL          1.50  
#>  8 S2          MNOP         NA     
#>  9 S3          ABCD         -5.94  
#> 10 S3          EFGH         10.4   
#> 11 S3          IJKL        -14.3   
#> 12 S3          MNOP         -6.87

由reprex package (v0.2.0) 于 2018 年 5 月 10 日创建。

【讨论】:

    猜你喜欢
    • 2020-11-14
    • 2020-08-28
    • 1970-01-01
    • 1970-01-01
    • 2021-08-13
    • 2020-03-08
    • 2020-08-22
    • 2016-06-11
    • 2020-06-17
    相关资源
    最近更新 更多