【发布时间】:2017-02-23 15:02:36
【问题描述】:
我有一个数据集,其简短版本如下所示:
> df
V1 V2
MID_R 1.243879014
MID 2.238147196
MID_Rcon 0.586581997
MID_U 0.833624164
MID -0.681462038
MID -0.593624936
MID_con 0.060862707
MID_con -0.764524044
MID_R -0.128464132
我已经编写了一个代码来只选择 MID 行并计算它们的方法:
MID_match <- c("MID^") # choosing specific pattern to search through conditions
MID <- df[grepl(paste(MID_match, collapse="|"), df$V1), ] # grouping across this pattern
MID$V2 <- as.numeric(as.character(MID$V2))
mean_MID <- mean(MID$V2) # calculating mean
MID_mean = rbind(MID_mean, data.frame(mean_MID))
我想要的第一行和第二行的输出应该是这样的:
> MID_match
[1] "MID"
> MID
V1 V2
MID 2.238147196
MID -0.681462038
MID -0.593624936
但是,我得到了包含字符串 MID 的所有行,例如,初始数据集:
> MID
V1 V2
MID_R 1.243879014
MID 2.238147196
MID_Rcon 0.586581997
MID_U 0.833624164
MID -0.681462038
MID -0.593624936
MID_con 0.060862707
MID_con -0.764524044
MID_R -0.128464132
我尝试使用 grep 函数,但它不起作用:
MID_match <- df$V1(grep("\\bMID\\b", df$V1))
关于如何提取确切的 MID 值的任何想法?
【问题讨论】: