【发布时间】:2020-01-29 14:52:50
【问题描述】:
我的data.table 包含每小时对引擎产生的功率的观察 (output) 和系统状态描述符 tag,它告诉引擎的所有组件都已打开。
数据
structure(list(time = structure(c(1517245200, 1517247000, 1517248800,
1517250600, 1517252400, 1517254200, 1517256000, 1517257800, 1517259600,
1517261400, 1517263200, 1517265000, 1517266800, 1517268600, 1517270400,
1517272200, 1517274000, 1517275800, 1517277600, 1517279400, 1517281200,
1517283000, 1517284800, 1517286600), class = c("POSIXct", "POSIXt"
), tzone = ""), output1 = c(160.03310020928, 159.706274495615,
159.803834736236, 159.753928429527, 159.54807802046, 159.21298848298,
158.904290018581, 158.683643772917, 158.670475839199, 158.793901799427,
158.886487460894, 159.167829223303, 159.66751884913, 159.1288534448,
159.141463186901, 160.116892086363, 160.517879769862, 160.615925580417,
160.915687799509, 161.590897854561, 161.568455821241, 161.411642091721,
161.811137570257, 162.193040254917), tag1 = c("evap only", "evap only",
"fog & evap", "fog & evap", "evap only", "evap only", "evap only",
"neither fog nor evap", "neither fog nor evap", "fog & evap", "evap only", "evap only",
"evap only", "fog & evap", "evap only", "fog & evap", "evap only",
"evap only", "evap only", "evap only", "fog & evap", "fog & evap",
"bad data", "neither fog nor evap")), row.names = c(NA, -24L
), class = c("data.table", "data.frame"))
您还可以使用以下方法生成一些示例数据:
sample_data <- data.table(time = seq.POSIXt(from = Sys.time(), by = 60*60*3, length.out = 100),
output = runif(n = 100, min = 130, max = 172),
tag = sample(x = c('evap only', 'bad data', 'neither fog nor evap', 'fog and evap'),
size = 100, replace = T))
我想按天分组(上面的示例数据只有两天,但实际数据有 3 年的数据)并找到每个 tag 对应的平均功率。我希望输出类似于:
time evap only fog & evap neither fog nor evap bad data
1: 2018-01-29 159.8391 160.0825 159.8491 161.8111
我尝试了以下代码,但结果不是我想要的形式。我正在使用.SDcols,因为实际数据集有大量其他列。
sample_data[, lapply(.SD, function(z){mean(z, na.rm = T)}), .SDcols = c('output1'), by = .(round_date(time, 'day'), tag1)]
round_date tag1 output1
1: 2018-01-30 evap only 159.8391
2: 2018-01-30 fog & evap 160.0825
3: 2018-01-30 neither fog nor evap 159.8491
4: 2018-01-30 bad data 161.8111
我已经看到以下有关堆栈溢出的问题。
- Create new data.table columns based on other columns
- Loop through data.table and create new columns basis some condition
- R data.table create new columns with standard names
- Add new columns to a data.table containing many variables
- Add multiple columns to R data.table in one function call?
- Assign multiple columns using := in data.table, by group
- Dynamically create new columns in data.table
- Creating new columns in data.table
有没有data.table 的方法来实现这一点?
【问题讨论】:
-
你在找
dcast(DT[, mean(output1), .(d=as.Date(time), tag1)], d ~ tag1, value.var="V1")吗?由于您想要的输出只有 1 个日期,因此很难说出您在寻找什么 -
如果您已经有了按日期计算的平均值,这不只是一个重塑问题吗? stackoverflow.com/questions/5890584/…
-
@chinsoon12 来自其他日期的数据最终将作为输出中的其他行。我添加了一个部分来生成一些带有附加日期的随机数据。
-
@RonakShah 我意识到我错过了什么,它现在可以工作,但我想知道是否有一种 data.table 方法可以实现这一点。
-
样本数据集的期望输出是什么
标签: r data.table