【问题标题】:how to use map function in r to find the range and quantile如何在 r 中使用 map 函数来查找范围和分位数
【发布时间】:2019-03-02 09:12:40
【问题描述】:

我首先在正态分布中模拟了 500 个大小为 55 的样本。

samples <- replicate(500, rnorm(55,mean=50, sd=10), simplify = FALSE)

1) 对于每个样本,我想要平均值、中位数、范围和第三四分位数。然后我需要将它们一起存储在一个数据框中。

这就是我所拥有的。我不确定范围或分位数。我尝试了 sapply 和 lapply 但不确定它们是如何工作的。

stats <- data.frame(
means = map_dbl(samples,mean),
medians = map_dbl(samples,median),
sd= map_dbl(samples,sd),

range= map_int(samples, max-min),
third_quantile=sapply(samples,quantile,type=3)
)

2) 然后绘制均值的采样分布(直方图)。 我尝试绘图,但我不知道如何获得平均值

stats <- gather(stats, key = "Trials", value = "Mean")

ggplot(stats,aes(x=Trials))+geom_histogram()

3) 然后我想在单个绘图窗口的(三个单独的图表)中绘制其他三个统计数据。

我知道我需要使用诸如收集和 facet_wrap 之类的东西,但我不知道该怎么做。

【问题讨论】:

    标签: r plot range tidyverse map-function


    【解决方案1】:

    你快到了。只需在有错误的地方定义匿名函数。

    library(tidyverse)
    
    set.seed(1234)    # Make the results reproducible
    
    samples <- replicate(500, rnorm(55,mean=50, sd=10), simplify = FALSE)
    str(samples)
    
    stats <- data.frame(
      means = map_dbl(samples, mean),
      medians = map_dbl(samples, median),
      sd = map_dbl(samples, sd),
      range = map_dbl(samples, function(x) diff(range(x))),
      third_quantile = map_dbl(samples, function(x) quantile(x, probs = 3/4, type = 3))
    )
    
    str(stats)
    #'data.frame':  500 obs. of  5 variables:
    # $ means         : num  49.8 51.5 52.2 50.2 51.6 ...
    # $ medians       : num  51.5 51.7 51 51.1 50.5 ...
    # $ sd            : num  9.55 7.81 11.43 8.97 10.75 ...
    # $ range         : num  38.5 37.2 54 36.7 60.2 ...
    # $ third_quantile: num  57.7 56.2 58.8 55.6 57 ...
    

    【讨论】:

      【解决方案2】:

      您正在使用的map_dbl 函数绝对不错,但是如果您最终还是想获得一个数据框,那么您可能更容易在开始时将列表转换为数据框,然后利用一些dplyr 函数。

      我首先映射列表,创建小标题,并将其与添加的 ID 绑定在一起。转换会创建一个包含样本值的列value。 summarise_at 允许您获取函数列表——在列表中提供名称会在结果数据框中设置名称。您可以使用purrr 的~. 表示法在需要的地方内联定义这些函数。减少map_dbl 等的次数。

      library(tidyverse)
      stats <- samples %>%
        map_dfr(as_tibble, .id = "sample") %>%
        group_by(sample) %>%
        summarise_at(vars(value), 
                     .funs = list(mean = mean, median = median, sd = sd,
                                  range = ~(max(.) - min(.)),
                                  third_quartile = ~quantile(., probs = 0.75)))
      
      head(stats)
      #> # A tibble: 6 x 6
      #>   sample  mean median    sd range third_quartile
      #>   <chr>  <dbl>  <dbl> <dbl> <dbl>          <dbl>
      #> 1 1       45.0   44.4  8.71  47.6           48.6
      #> 2 10      51.0   52.0  9.55  49.3           56.2
      #> 3 100     51.6   52.2 10.4   60.7           58.1
      #> 4 101     51.6   51.1  9.92  37.6           57.2
      #> 5 102     49.1   48.2  9.65  39.8           57.0
      #> 6 103     52.2   51.3 10.1   47.4           58.5
      

      接下来,在您的代码中 gathered 数据——这通常是人们在 SO 上需要的解决方案——但如果你只是想显示平均列,你可以按原样使用它。

      ggplot(stats, aes(x = mean)) +
        geom_histogram()
      

      【讨论】:

        猜你喜欢
        • 2015-10-04
        • 1970-01-01
        • 1970-01-01
        • 2017-12-19
        • 2015-03-04
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多