【问题标题】:adding error bars and significance bars to binary data ggplot向二进制数据ggplot添加误差线和重要性条
【发布时间】:2021-05-15 02:06:48
【问题描述】:

我有以下数据集:

df <- data.frame(replicate(2,sample(0:1,30,rep=TRUE)))
df <- reshape(data=df, varying=list(1:2), 
        direction="long", 
        times = names(df), 
        timevar="Type",
        v.names="Score")

并创建我使用的条形图

ggplot(df ,aes(x=Score, fill = Type))+
  geom_bar(position="dodge", color="black")

这给了我们

我想添加类似于

的误差线

对于分级数据,我只会使用 add = "mean_ci",但将其添加到此函数中没有任何作用

ggplot(df ,aes(x=Score, fill = Type), add = "mean_ci")+
  geom_bar(position="dodge", color="black")

如何为我的二进制数据条形图获得漂亮的误差线?

此外,我还想添加重要性条,因此条形图最终看起来类似于以下内容

但是我不知道如何使用看起来像我上面的数据集。

非常感谢任何帮助。

【问题讨论】:

    标签: r ggplot2 errorbar


    【解决方案1】:

    从原始 df 中,只需使用 table 创建另一个数据框,然后从那里计算 count 和 sd。二项式变量,因此在这里使用定义为 np(1-p) 的方差:

    df <- data.frame(replicate(2,sample(0:1,30,rep=TRUE)))
    df1 <- reshape(data=df, varying=list(1:2), 
                  direction="long", 
                  times = names(df), 
                  timevar="Type",
                  v.names="Score")
    
    t1 <- as.data.frame(table(df$X1))
    t1$Type <- "X1"
    t1$sd <- sqrt(sum(t1$Freq)*t1$Freq/sum(t1$Freq)*(1-t1$Freq/sum(t1$Freq)))
    t2 <- as.data.frame(table(df$X2))
    t2$Type <- "X2"
    t2$sd <- sqrt(sum(t2$Freq)*t2$Freq/sum(t2$Freq)*(1-t2$Freq/sum(t2$Freq)))
    df_tbl <- rbind(t1,t2)
    names(df_tbl) <- c("Score","count","Type","sd")
    df_tbl$marg <- df_tbl$count + df_tbl$sd + 1
                                
    y0 <- max(df_tbl[which(df_tbl$Score==0),'marg'])
    y1 <- max(df_tbl[which(df_tbl$Score==1),'marg'])
    prob0 <- prop.test(c(nrow(df)-sum(df$X1), nrow(df)-sum(df$X2)), c(nrow(df), nrow(df)))
    prob1 <- prop.test(c(sum(df$X1), sum(df$X2)), c(nrow(df), nrow(df)))
    prob <- c(prob0$p.value, prob1$p.value)
    prob_sig <- as.character(cut(prob,breaks=c(-1,0.0005,0.001,0.01,max(prob)),
                           labels=c("***","**","*","NS")))
    
    ggplot(df_tbl, aes(x=Score, y=count, fill=Type)) + 
      geom_bar(stat="identity", color="black", 
               position=position_dodge()) +
      geom_errorbar(aes(ymin=count-sd, ymax=count+sd), width=.2,
                    position=position_dodge(.9)) +
      ylim(c(0,max(df_tbl$count+df_tbl$sd)+4)) +
      geom_segment(aes(x = 0.8, xend = 1.2, y = y0, yend = y0)) +
      geom_segment(aes(x = 1.8, xend = 2.2, y = y1, yend = y1)) +
      geom_text(aes(x = 1, y = y0 + 0.5, label=prob_sig[1])) +
      geom_text(aes(x = 2, y = y1 + 0.5, label=prob_sig[2]))
    

    【讨论】:

    • 感谢您提供有用的回答,但是我不太明白“df_tbl$marg”之后发生的事情以及有关重要性条的所有内容 - 我正在寻找一种比较比例的方法x1 和 x2 0,以及 x1 和 x2 1,看看它们是否有显着差异(类似于 stat_compare_means,但适用于二进制数据。理想情况下,还带有末端带有“钩子”的实心重要性条。
    • @Icewaffle 为此,您必须执行 prop.test。我更新了代码以使用正确的概率值。由于您正在进行组间比较,因此您不能完全为那些做错误栏(带有“钩子”)。要么通过另一个 geom_segment 手动插入这些钩子,要么选择 lineshape 等,以便它们可以区分。希望对您有所帮助!
    • 谢谢,很有用!最后一个问题——这些误差线代表标准偏差还是置信区间?我忘了指定它,但我正在寻找置信区间。
    • @Icewaffle 我使用了标准。偏差,但要获得 95% 的置信区间,您可以在 geom_errorbar 中执行 ymin=count-1.96*sd/sqrt(nrow(df)), ymax=count+1.96*sd/sqrt(nrow(df))。
    猜你喜欢
    • 2018-06-24
    • 2017-06-12
    • 2016-01-03
    • 1970-01-01
    • 1970-01-01
    • 2019-12-04
    • 1970-01-01
    • 1970-01-01
    • 2020-03-25
    相关资源
    最近更新 更多