【问题标题】:count number of values in colum that are in a certain interval, by subgroup按子组计算列中某个区间内的值的数量
【发布时间】:2020-01-04 15:20:50
【问题描述】:

我在下面的代码中做错了,因为列选择似乎不是按子组。想不通。

library(data.table)
time <- c(1,1,1,1,1,1,1,2,2,2,2,2,2,2,2,3)
mass <- c(2,2,2,2,2,2,2,5,5.1,4.9,5,5,5,10,10,10)
expected_check <- c(7,7,7,7,7,7,7,6,6,6,6,6,6,2,2,1)
dt <- data.table(time, mass, expected_check)

dt[, check:= sapply(mass, function(i) length(dt[,mass] <= (i+0.5) & dt[,mass] >= (i-0.5))), by = time]

我想同时检查每个质量值,在定义的时间间隔内质量列中有多少个值。我认为 dt[,mass] 选择整个列并且不跟随最终的 by = time。 我尝试过的所有东西都遇到了同样的问题。

【问题讨论】:

  • 你能检查一下你的expected_check,因为它没有在'时间'中关注组
  • 我的意思是第 14 次和第 15 次观察在时间 2 也属于同一组

标签: r data.table


【解决方案1】:

在 OP 的代码中,我们需要 .SD[["mass"]] 而不是选择整个列的 dt[, mass] 来确保选择分组内的值

dt[, sapply(mass, function(i) length(.SD[["mass"]] <= (i+0.5) & 
           .SD[["mass"]] >= (i-0.5))), by = time]

【讨论】:

  • 我认为没有 .SD 也能正常工作,即dt[, check:= sapply(mass, function(i) length(mass &lt;= (i+0.5) &amp; mass &gt;= (i-0.5))), by = time]
【解决方案2】:

产生预期结果的另一种方法是使用非等连接:

time <- c(1,1,1,1,1,1,1,2,2,2,2,2,2,2,2,3)
mass <- c(2,2,2,2,2,2,2,5,5.1,4.9,5,5,5,10,10,10)
expected_check <- c(7,7,7,7,7,7,7,6,6,6,6,6,6,2,2,1)

dt <- data.table(time, mass, expected_check)

dt[dt[ , .(time, high_mass = mass + 0.5, low_mass = mass - 0.5)],
   on = .(time,
         mass <= high_mass,
         mass >= low_mass),
   check := .N,
   by = .EACHI]

dt

关于另一种方法,我认为您想使用sum() 而不是length()。使用@chinsoon12 的代码:

dt[, wrong_check := sapply(mass, function(i) length(mass <= (i+0.5) & mass >= (i-0.5))), by = time]
dt[, correct_check := sapply(mass, function(i) sum(mass <= (i+0.5) & mass >= (i-0.5))), by = time]
dt

     time  mass expected_check check wrong_check correct_check
    <num> <num>          <num> <int>       <int>         <int>
 1:     1   2.0              7     7           7             7
 2:     1   2.0              7     7           7             7
 3:     1   2.0              7     7           7             7
 4:     1   2.0              7     7           7             7
 5:     1   2.0              7     7           7             7
 6:     1   2.0              7     7           7             7
 7:     1   2.0              7     7           7             7
 8:     2   5.0              6     6           8             6
 9:     2   5.1              6     6           8             6
10:     2   4.9              6     6           8             6
11:     2   5.0              6     6           8             6
12:     2   5.0              6     6           8             6
13:     2   5.0              6     6           8             6
14:     2  10.0              2     2           8             2
15:     2  10.0              2     2           8             2
16:     3  10.0              1     1           1             1

问题在于 length() 返回向量的结果 - c(T, F, F, F) 与 c(T, T, T, T) 具有相同的长度。相反,您的 expected_check 与 TRUE 结果的数量相同。我们可以通过使用sum() 来获得它。

【讨论】:

    猜你喜欢
    • 2020-04-03
    • 2014-12-17
    • 2012-03-22
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-03-24
    相关资源
    最近更新 更多