【问题标题】:Adding error bars on multiple data series using ggplot使用 ggplot 在多个数据系列上添加误差线
【发布时间】:2017-06-05 03:05:21
【问题描述】:

我记录了 5 次治疗(包括对照)在几个月内的一些数据。我正在使用 ggplot 将数据绘制为时间序列,并为每个日期生成了原始数据均值和标准误差的数据框。

我正在尝试在同一张图表上绘制所有五种治疗方法并用它显示误差线。我能够a)。绘制一个治疗组并显示误差线和 b)。绘制所有五种处理方法,但不显示误差线。

这是我的数据(我只包括了两种处理方法来保持这里的整洁)

       dates   c_mean_am  c_se_am    T1_mean_am  T1_se_am  
1 2017-01-31   284.135   27.43111     228.935     23.39037    
2 2017-02-09   226.944   13.08237     173.241     13.42946    
3 2017-02-23   281.135   15.89709     252.665     20.73417   
4 2017-03-14   265.655   15.29930     238.225     17.47501 
5 2017-04-06   312.785   13.08237     237.485     13.42946 
  • c_mean_am = 控制手段
  • c_se_am = 控件的标准错误
  • T1_mean_am = 治疗 1 意味着
  • T1_se_am = 治疗 1 的标准误差

这是我实现上述选项 a) 的代码

ggplot(summary, aes(x=dates, y=c_mean_am),xlab="Date") + 
    geom_point(shape = 19, size = 2,color="blue") + 
    geom_line(color="blue") + 
    geom_errorbar(aes(x=dates, ymin=c_mean_am-c_se_am, ymax=c_mean_am+c_se_am), color="blue", width=0.25) 

这是上面选项 b) 的代码

sp <- ggplot(summary,aes(dates,y = Cond,color=Treatment)) + 
    geom_line(aes(y = c_mean_am, color = "Control")) + 
    geom_line(aes(y = T1_mean_am, color = "T1")) + 
    geom_point(aes(y = c_mean_am, color = "Control")) + 
    geom_point(aes(y = T1_mean_am, color = "T1"))

sp2<- sp + 
    scale_color_manual(breaks = c("Control", "T1","T2"), values=c("blue", "yellow"))

sp2

如何使用与点和线相同的颜色来获得第二个图上的误差线?

谢谢

AB

【问题讨论】:

    标签: r ggplot2


    【解决方案1】:

    首先将您的数据转换为长格式:

    df <- df %>% 
     gather(mean_type, mean_val, c_mean_am, T1_mean_am) %>% 
     gather(se_type, se_val, c_se_am, T1_se_am)
    
    
    ggplot(df, aes(dates, mean_val, colour=mean_type)) + 
        geom_line() + 
        geom_point() + 
        geom_errorbar(aes(ymin=mean_val-se_val, ymax=mean_val+se_val))
    

    编辑:tidyr 操作的解释

    new.dat <- mtcars %>%  # taking mtcars as the starting data.frame
            select(gear, cyl, mpg, qsec) %>% 
              # equivalent to mtcars[, c("gear", "cyl", "mpg", "qsec")]; to simplify the example
            gather(key=type, value=val, gear, cyl) %>% 
              # convert the data into a long form with 64 rows, with new factor column "type" and numeric column "val". "gear" and "cyl" are removed while "mpg" and "qsec" remain
    
    new.dat[c(1:3, 33:35),]
    
    #     mpg  qsec type val
    # 1  21.0 16.46 gear   4
    # 2  21.0 17.02 gear   4
    # 3  22.8 18.61 gear   4
    # 33 21.0 16.46  cyl   6
    # 34 21.0 17.02  cyl   6
    # 35 22.8 18.61  cyl   4
    

    对于长形式的数据,您可以使用新的标识符形式(“类型”)进行绘图,例如

    ggplot(new.dat, aes(val, mpg, fill=type)) + 
       geom_col(position="dodge")
    

    长格式也可用于在不同方面进行绘图,例如

    ggplot(new.dat, aes(val, mpg, colour=type)) + 
        geom_point() + 
        facet_wrap(~type) 
    

    【讨论】:

    • 感谢您的快速回复。我使用了您的代码并收到以下错误消息“错误:is.character(x) is not TRUE”我做错了什么?
    • 你的data.frame的结构是什么?错误发生在哪一步:数据操作还是绘图?
    • 我对 R 很陌生,所以你可能需要指导我完成这个。数据框的结构是什么意思?我基本上只是复制了您的代码并将其放在我的数据操作之后以生成绘图并得到错误。我需要安装“tidyr”和“devtools”包。
    • 对不起。我假设你知道dplyr 和tidyr,因为你在ggplot2 上运行。无论如何,只需执行install.packages("tidyverse") 来下载/安装所需的包,然后运行library(tidyverse) 来加载功能所需的包。至于结构,如果你输入str(df),你会得到每列的类摘要(例如str(iris))
    • 所以我添加了软件包,但仍然收到相同的错误消息。我根本无法生成“df”来查找结构。我只是将您的代码直接复制并粘贴到我的代码之后,我曾经用两个数据系列绘制图表。我正在使用我自己的流程数据的数据框,我这里的数据只是其中的一部分。我可以在上面使用 Str 。除了日期列是日期之外,所有列都是“num”。
    【解决方案2】:

    接受的答案似乎包含数据的方式错误gathered(又名pivot_longer in packageVersion("tidyr") &gt;= 1.0.0)重复每个点和错误栏。误差线很明显,但如果您将geom_point() 替换为geom_jitter(),您将看到与两个误差线相对应的两个点。这给其他人造成了一些confusion,所以我想为后代提供一个更正的解决方案。

    这是避免这种重复的另一种方法:

    # load necessary packages
    library(tidyverse)
    
    # create data from question
    df <-
      structure(
        list(
          dates = c(
            "2017-01-31",
            "2017-02-09",
            "2017-02-23",
            "2017-03-14",
            "2017-04-06"
          ),
          c_mean_am = c(284.135, 226.944,
                        281.135, 265.655, 312.785),
          c_se_am = c(27.43111, 13.08237, 15.89709,
                      15.2993, 13.08237),
          T1_mean_am = c(228.935, 173.241, 252.665,
                         238.225, 237.485),
          T1_se_am = c(23.39037, 13.42946, 20.73417,
                       17.47501, 13.42946)
        ),
        class = "data.frame",
        row.names = c("1",
                      "2", "3", "4", "5")
      )
    
    # pivot df long and confirm that there's only one value per group per timepoint
    df_long <- df %>%
      pivot_longer(
        cols = -dates,
        names_to = c("treatment_group", ".value"),
        names_pattern = "(.*)_(.*_am)"
      ) 
    
    df_long
    
    # # A tibble: 10 x 4
    #    dates      treatment_group mean_am se_am
    #    <chr>      <chr>             <dbl> <dbl>
    #  1 2017-01-31 c                  284.  27.4
    #  2 2017-01-31 T1                 229.  23.4
    #  3 2017-02-09 c                  227.  13.1
    #  4 2017-02-09 T1                 173.  13.4
    #  5 2017-02-23 c                  281.  15.9
    #  6 2017-02-23 T1                 253.  20.7
    #  7 2017-03-14 c                  266.  15.3
    #  8 2017-03-14 T1                 238.  17.5
    #  9 2017-04-06 c                  313.  13.1
    # 10 2017-04-06 T1                 237.  13.4
    

    现在您可以在每个时间点为每个组绘制并获得只有一个误差线和一个点的预期图表。

    df_long %>%
      ggplot(aes(x = dates, y = mean_am, colour = treatment_group)) + 
      geom_line(aes(group = treatment_group)) + 
      geom_point() + 
      geom_errorbar(aes(ymin = mean_am - se_am, ymax = mean_am + se_am))
    

    这会产生这个情节:

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-01-03
      • 2020-03-25
      • 2018-06-24
      • 1970-01-01
      • 1970-01-01
      • 2021-10-07
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多