【问题标题】:Creating a variable conditional on the sum of a variable by group以组变量的总和为条件创建变量
【发布时间】:2020-07-14 12:58:33
【问题描述】:

我有一个data.table,如下:

panelID = c(1:50)   
year    = c(2001:2010)
country = c("NLD", "BEL", "GER")
urban   = c("A", "B", "C")
indust  = c("D", "E", "F")
sizes   = c(1, 2, 3, 4, 5)
n <- 2

library(data.table)

set.seed(123)
DT <- data.table(
    panelID = rep(sample(panelID), each = n),
    country = rep(sample(country, length(panelID), replace = T), each = n),
    year    = c(replicate(length(panelID), sample(year, n))),
    some_NA = sample(0:5, 6),                                             
    some_NA_factor = sample(0:5, 6),
    industry       = rep(sample(indust, length(panelID), replace = T), each = n),
    urbanisation   = rep(sample(urban, length(panelID), replace = T), each = n),
    size      = rep(sample(sizes, length(panelID), replace = T), each = n),
    norm      = round(runif(100)/10, 2),
    sales     = round(rnorm(10, 10, 10), 2),
    Happiness = sample(10, 10),
    Sex       = round(rnorm(10, 0.75, 0.3), 2),
    Age       = sample(100, 100),
    Educ      = round(rnorm(10, 0.75, 0.3), 2)
)        
DT [, uniqueID := .I]  # Creates a unique ID     
DT[DT == 0] <- NA 
DT$sales[DT$sales< 0] <- NA 
DT <- as.data.frame(DT)

我想要的是panelIDs 的数量,其中size 的总和等于8。所以我想我会这样做:

DT[sum(size)==8, condition:=1, by=panelID]

我在这里做错了什么?

【问题讨论】:

  • 你需要这个吗:DT[,conditional:=ifelse(sum(size)==8,1,0),by=panelID][]?
  • 好的,谢谢!比我想象的要复杂..

标签: r sum data.table conditional-statements


【解决方案1】:

与data.table:

DT[,conditional:=ifelse(sum(size)==8,1,0),by=panelID][]
# To get the lengths of those which are True(1), save the above as res
#nrow(res[res[,conditional==1],"panelID"])

或者就像@chinsoon12 建议的那样:

DT[, conditional := +(sum(size)==8), panelID]

结果(头部):

 panelID country year some_NA some_NA_factor industry urbanisation size norm sales
1:      31     GER 2010       4              1        F            C    4 0.09  5.63
2:      31     GER 2005       2             NA        F            C    4 0.03 13.31
3:      15     NLD 2005      NA              4        D            C    3 0.05    NA
4:      15     NLD 2008       1              5        D            C    3 0.01 12.12
5:      14     BEL 2003       5              3        E            B    1 0.09 22.37
6:      14     BEL 2002       3              2        E            B    1 0.04 30.38
   Happiness  Sex Age Educ uniqueID conditional
1:         7 0.69  62 0.25        1           1
2:         3 1.00  10 1.31        2           1
3:        10 0.66  59 0.73        3           0
4:         9 0.85  49 0.88        4           0
5:         2 0.34   7 0.90        5           0
6:         5 0.84  61 1.11        6           0

【讨论】:

    【解决方案2】:

    你可以用 dplyr 来做

    你可以通过使用dplyr的这段代码来实现你想要的:

    library(dplyr)
    DT %>%
      group_by(panelID) %>%
      summarize(sum = sum(size)) %>%
      filter(sum == 8) %>%
      pull(panelID)
    
    #Output
    [1] 11 14 15 16 18 27 28 34 38 45
    

    编辑

    如果要显示面板数量,可以将pull(panelID)改成count(),或者在末尾加上lenght(),如下所示:

    library(dplyr)
    DT %>%
      group_by(panelID) %>%
      summarize(sum = sum(size)) %>%
      filter(sum == 8) %>%
      pull(panelID) %>%
      length()
    
    #Output
    [1] 10
    

    希望这会有所帮助。

    【讨论】:

    • 感谢您的回答,路易斯!如果我只想要panelIDs 的数量,我会用什么代替pull?我试过count 和sum 但这不起作用..
    • 不客气!如果您想要面板的数量,只需删除 pull(panelID) 并使用 count() 更改它。
    【解决方案3】:

    我刚刚删除了as.data.frame()。我使用连接将size 的总和与panelID 正确对齐。

    我不明白的是,如果您想要满足总和给出的条件的panelID 的值,我猜是panelID。或者,如果您想要有多少 panelID(即个人?)满足条件。

    在前一种情况下,你可以这样做:

    panelID = c(1:50)   
    year    = c(2001:2010)
    country = c("NLD", "BEL", "GER")
    urban   = c("A", "B", "C")
    indust  = c("D", "E", "F")
    sizes   = c(1, 2, 3, 4, 5)
    n <- 2
    
    library(data.table)
    
    set.seed(123)
    DT <- data.table(
      panelID = rep(sample(panelID), each = n),
      country = rep(sample(country, length(panelID), replace = T), each = n),
      year    = c(replicate(length(panelID), sample(year, n))),
      some_NA = sample(0:5, 6),                                             
      some_NA_factor = sample(0:5, 6),
      industry       = rep(sample(indust, length(panelID), replace = T), each = n),
      urbanisation   = rep(sample(urban, length(panelID), replace = T), each = n),
      size      = rep(sample(sizes, length(panelID), replace = T), each = n),
      norm      = round(runif(100)/10, 2),
      sales     = round(rnorm(10, 10, 10), 2),
      Happiness = sample(10, 10),
      Sex       = round(rnorm(10, 0.75, 0.3), 2),
      Age       = sample(100, 100),
      Educ      = round(rnorm(10, 0.75, 0.3), 2)
    )        
    DT [, uniqueID := .I]  # Creates a unique ID     
    DT[DT == 0] <- NA 
    DT$sales[DT$sales< 0] <- NA 
    
    dt_sum = DT[ , .(size_sum = sum(size) ), by = panelID ]
    setkey( dt_sum, panelID )
    setkey( DT, panelID )
    
    DT = DT[ dt_sum ]
    final = DT[ size_sum == 8, .N, by = panelID ]
    > final
        panelID N
     1:       6 2
     2:       8 2
     3:       9 2
     4:      11 2
     5:      18 2
     6:      22 2
     7:      28 2
     8:      30 2
     9:      31 2
    10:      38 2
    

    后一种情况,你只需统计final的行数:

    > nrow( final )
    6
    

    【讨论】:

    • 感谢您的回答!我很遗憾地寻找后一种情况;)请参阅 NelsonGon 的回答。
    猜你喜欢
    • 1970-01-01
    • 2015-05-25
    • 2017-10-30
    • 1970-01-01
    • 2021-04-02
    • 1970-01-01
    • 2016-10-21
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多