【问题标题】:Extracting columns with constant numbers in R data.frames在 R data.frames 中提取具有常数的列
【发布时间】:2020-02-07 08:58:45
【问题描述】:

在 data.frame DATA 中,我有一些列是第一列的唯一行中的常数,称为 study.name。例如,对于Shin.Ellis 的所有行,列ESL 和prof 是constant,对于Trus.Hsu 的所有行,列是constant 等等。包括Shin.Ellis 和Trus.Hsu,共有8 个唯一的study.name 行。

但是在下面我的split.default() 调用之后,对于这样的常量,我如何才能为唯一的study.name 下的所有行获取一个数据点(例如,一个用于Shin.Ellis,一个用于Trus.Hsu 等)变量? (即总共 8 行)

例如,在我的split.default() 之后,所有名为ESL 的变量都显示只有8 行,每行对应一个唯一的study.name。

我想要的输出 ONLY ESL 和 prof 如下所示。

注意:这是玩具数据。我们首先应该找到常量变量。非常感谢功能性答案。

DATA <- read.csv("https://raw.githubusercontent.com/izeh/m/master/irr.csv", h = T)[-(2:3)]
DATA <- setNames(DATA, sub("\\.\\d+$", "", names(DATA)))

tbl <- table(names(DATA))
nm2 <- names(which(tbl==max(tbl)))

L <- split.default(DATA[names(DATA) %in% nm2], names(DATA)[names(DATA) %in% nm2])


## FIRST 8 ROWS of `DATA`:

#    study.name ESL prof scope type   ESL   prof   scope   type
# 1  Shin.Ellis   1    2     1    1     1      2       1      1
# 2  Shin.Ellis   1    2     1    1     1      2       1      1
# 3  Shin.Ellis   1    2     1    2     1      2       1      1
# 4  Shin.Ellis   1    2     1    2     1      2       1      1
# 5  Shin.Ellis   1    2    NA   NA     1      2      NA     NA
# 6  Shin.Ellis   1    2    NA   NA     1      2      NA     NA
# 7    Trus.Hsu   2    2     2    1     2      2       1      1
# 8    Trus.Hsu   2    2    NA   NA     2      2      NA     NA
# .     ...       .    .     .    .     .      .       .      . # `DATA` has 54 rows overall

ESL 和 prof 在split.default() 调用后的所需输出:

# $ESL            ## 8 unique rows for 8 unique `study.name`
#    ESL ESL.1
# 1    1     1
# 7    2     2
# 9    1     1
# 17   1     1
# 23   1     1
# 35   1     1
# 37   2     2
# 49   2     2


# $prof           ## 8 unique rows for 8 unique `study.name`
#    prof prof.1
# 1     2      2
# 7     2      2
# 9     3      3
# 17    2      2
# 23    2      2
# 35    2      2
# 37   NA     NA
# 49    2      2

【问题讨论】:

    标签: r list function loops dataframe


    【解决方案1】:

    我们可以先找到常量列,然后用lapply循环遍历它们,在每个study.name中只选择它们的第一行。

    is_constant <- function(x) length(unique(x)) == 1L 
    cols <- names(Filter(all, aggregate(.~study.name, DATA, is_constant)[-1]))
    
    L[cols] <- lapply(L[cols], function(x) 
                          x[ave(x[[1]], DATA$study.name, FUN = seq_along) == 1, ])
    L
    
    #$ESL
    #   ESL ESL.1
    #1    1     1
    #7    2     2
    #9    1     1
    #17   1     1
    #23   1     1
    #35   1     1
    #37   2     2
    #49   2     2
    
    #$prof
    #   prof prof.1
    #1     2      2
    #7     2      2
    #9     3      3
    #17    2      2
    #23    2      2
    #35    2      2
    #37   NA     NA
    #49    2      2
    #.....
    

    【讨论】:

    • @rnorouzian 我认为在这种情况下,您可以将内部函数更改为 x[ave(seq_along(x[[1]]), DATA$study.name, FUN = seq_along) == 1, ])
    • 嗨,Ronak,@rnorouzian 的问题对您来说有意义还是他遗漏了什么? -- 谢谢。
    • 嗨...如果您还有其他问题,请提出新问题。
    • 亲爱的 Ronak 你能看到the question 和问题吗?
    • 我再次在您的解决方案中发现了一个错误(您在上面的评论中编辑的那个)。要查看问题,假设DATA 是:a &lt;- data.frame(study.name = c(1,1,2,3), mod.s=c(3,3,1,2), mod.g=c(1,1,3,1)); b &lt;- data.frame(study.name = c(1,1,2,3), mod.s=c(3,3,2,2), mod.g=c(1,2,3,2)); DATA &lt;- cbind(a,b),同样L 是:L &lt;- split(DATA, DATA$study.name)。现在,如果您运行以cols 开头的代码不应返回任何内容,因为"mod.s" 和"mod.g" 在DATA 中不是常量,但它错误地返回"mod.s" 和"mod.g"?你能帮忙吗?
    【解决方案2】:

    我们可以使用aggregate创建预期的输出

    is_constant <- function(x) length(unique(x)) == 1L 
    nm1 <-  names(which(!colSums(!aggregate(.~ study.name, DATA, is_constant)[-1])))
    L[nm1] <- lapply(L[nm1], function(x) aggregate(x, 
       list(factor(DATA$study.name, levels = unique(DATA$study.name))), 
              FUN = head, 1)[-1])
    L
    #$ESL
    #  ESL ESL.1
    #1   1     1
    #2   2     2
    #3   1     1
    #4   1     1
    #5   1     1
    #6   1     1
    #7   2     2
    #8   2     2
    
    #$prof
    #  prof prof.1
    #1    2      2
    #2    2      2
    #3    3      3
    #4    2      2
    #5    2      2
    #6    2      2
    #7   NA     NA
    #8    2      2
    
    #$scope
    #...
    

    【讨论】:

    • 后续,假设 DATA 在我的原始帖子中如下:DATA &lt;- read.csv("https://raw.githubusercontent.com/izeh/m/master/c.csv")。根据我在我的问题中提出的问题,您的代码:nm1 应该返回 3 个名称:"setting" "prof" "random" 但是,它也错误地返回 "error"。但是,"error" 并不是所有study.names 的常数。比如看study.name == "Sun",你看到在一行"error"是2&在另一行"error"是0,但是为什么"error"是作为一个常量变量返回的呢?
    • @rnorouzian 你能发布一个新问题吗?谢谢
    猜你喜欢
    • 2020-02-13
    • 2017-06-10
    • 2017-12-12
    • 2019-12-29
    • 1970-01-01
    • 2019-11-10
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多