【问题标题】:Translating time stamps (start, end) into time series data. Errors with align.time() and colnames将时间戳(开始、结束)转换为时间序列数据。 align.time() 和 colnames 的错误
【发布时间】:2013-06-26 03:18:17
【问题描述】:

我是 R 新手,但是在学习了入门课程并玩了一点之后,我希望它可以 1) 更优雅地解决我的建模目标(与 Excel 相比,这是我的备用计划)和 2 ) 是从这个项目中带走的有用技能。

任务/目标:

我正在尝试使用驾驶日志数据来模拟和模拟电动汽车的潜在能量和温室气体排放。具体来说:

  1. 我有想要翻译成的驾驶日志数据(开始和结束时间戳,以及数千名司机的其他数据 - 基本示例如下):
  2. 24 小时时间序列数据,因此对于 24 小时期间的每一分钟,我确切地知道谁在驾驶车辆,以及它属于什么“行程”(对于那个司机)。我这里的问题集中在这个问题上。

我想要的输出类型: 注意:此输出与下面提供的示例数据相关。我以某天的前十分钟和一些理论旅行为例

对于这个问题不是必需的,但知道可能有用:我将使用上述输出来交叉引用其他特定于驾驶员的数据,以根据与相关的事物计算每分钟的汽油(或电力)消耗量该行程,例如停车位置或行程距离。我想在 R 中执行此操作,但在进行此步骤之前必须先弄清楚上述问题。

我目前的解决方案是基于:

问题:

简化数据示例:

a <- c("A","A","A","B","B","B","C","C","C")
b <- c(1, 2, 3, 1, 2, 3, 1, 2, 3)
c <- as.POSIXct(c(0.29167, 0.59375, 0.83333, 0.45833, 0.55347, 0.27083, 0.34375, 0.39236, 0.35417)*24*3600 + as.POSIXct("2013-1-1 00:00") )
d <- as.POSIXct(c(0.334027778, 0.614583333, 0.875, 0.461805556, 0.563888889, 0.295138889, 0.375, 0.503472222, 0.364583333)*24*3600 + as.POSIXct("2013-1-1 00:00"))
e <- c(2, 8, 2, 5, 5, 2, 5, 5, 2)
f <- as.POSIXct(c(0, 0.875, 0, 0.479166666666667, 0.580555555555556, 0.489583333333333, 0.430555555555556, 0.541666666666667, 0.711805555555555)*24*3600 + as.POSIXct("2013-1-1 00:00"))
g <- as.POSIXct(c(0, 0.885, 0, 0.482638888888889, 0.588194444444444, 0.496527777777778, 0.454861111111111, 0.559027777777778, 0.753472222222222)*24*3600 + as.POSIXct("2013-1-1 00:00"))
h <- c(0, 1, 0, 1, 4, 8, 8, 1, 5)
i <- as.POSIXct(c(0, 0, 0, 0.729166666666667, 0.595833333333333, 0.534722222222222, 0.59375, 0.779861111111111, 0.753472222222222)*24*3600 + as.POSIXct("2013-1-1 00:00"))
j <- as.POSIXct(c(0, 0, 0, 0.736111111111111, 0.605555555555556, 0.541666666666667, 0.611111111111111, 0.788194444444445, 0.75625)*24*3600 + as.POSIXct("2013-1-1 00:00"))
k <- c(0, 0, 0, 4, 4, 2, 5, 8,1)
testdata <- data.frame(a,b,c,d,e,f,g,h,i,j,k)
names(testdata) <- c("id", "Day", "trip1_start", "trip1_end", "trip1_purpose", "trip2_start", "trip2_end", "trip2_purpose", "trip3_start", "trip3_end", "trip3_purpose")

在这个示例数据中,我有三个司机(id = A、B、C),他们每个人在三个不同的日子(天 = 1、2、3)开车。请注意,某些司机可能有不同的行程次数。时间戳指示驾驶活动的开始和结束时间。

然后我为一整天(2013 年 1 月 1 日)创建分钟间隔

start.min <- as.POSIXct("2013-01-01 00:00:00 PST")
end.max <- as.POSIXct("2013-01-01 23:59:59 PST")
tinterval <- seq.POSIXt(start.min, end.max, na.rm=T, by = "mins")

在给定用户正在开车的几分钟内插入“1”:

out1 <- xts(,align.time(tinterval,60))
# loop over each user
for(i in 1:NROW(testdata)) {
  # paste the start / end times into an xts-style range
  timeRange <- paste(format(testdata[i,c("trip1_start","trip1_end")]),collapse="/")
  # add the minute "by parameter" for timeBasedSeq
  timeRange <- paste(timeRange,"M",sep="/")
  # create the by-minute sequence and align to minutes to match "out"
  timeSeq <- align.time(timeBasedSeq(timeRange),60)
  # create xts object with "1" entries for times between start and end
  temp1 <- xts(rep(1,length(timeSeq)),timeSeq)
  # merge temp1 with out and fill non-matching timestamps with "0"
  out1 <- merge(out1, temp1, fill=0)
}
# add column names
colnames(out1) <- paste(testdata[,1], testdata[,2], sep = ".")

我们的想法是为每次旅行重复此操作,例如out2、out3 等,其中我会用“2”、“3”等填充任何驾驶时段,然后汇总/合并所有生成的 outx 数据帧,最终得到所需结果。

不幸的是,当我尝试为 out2 重复此操作时...

out2 <- xts(,align.time(tinterval,60))
for(i in 1:NROW(testdata)) {
  timeRange2 <- paste(format(testdata[i,c("trip2_start","trip2_end")]),collapse="/")
  timeRange2 <- paste(timeRange2,"M",sep="/")
  timeSeq2 <- align.time(timeBasedSeq(timeRange2),60)
  temp2 <- xts(rep(2,length(timeSeq2)),timeSeq2)
  out2 <- merge(out2, temp2, fill=0)
}
colnames(out2) <- paste(testdata[,1], testdata[,2], sep = ".")
head(out2)

我收到以下错误:

  • UseMethod("align.time") 中的错误:没有适用于 'align.time' 的方法应用于“Date”类的对象
  • colnames&lt;-(*tmp*, value = c("A.1", "A.2", "A.3", "B.1", "B.2", 中的错误:尝试在具有较少的对象上设置“colnames” 超过二维

我的 out2 代码有什么问题?

我可以了解其他更好的解决方案或软件包吗?

我意识到这可能是一种非常迂回的方式来获得我想要的输出。

任何帮助将不胜感激。

【问题讨论】:

  • testdata的数据是怎么给你的?我问是因为如果数据是长格式会更简单。
  • 有一段旅程在开始之前就结束了。
  • @JoshuaUlrich:testdata 是我拥有的数据的一个非常简化的样本,但我将从中提取的基础知识。

标签: r time time-series xts


【解决方案1】:

这不是您问题的答案。老实说,我不清楚您在图像中显示的数据与您的数据示例之间的转换。看来您无法重现此数据。所以这里有一个函数来生成数据的可重现示例。我认为它至少可以用来验证您的模型。

数据采样函数

library(reshape2)
start.min <- as.POSIXct("2013-01-01 00:00:00 PST")
hours.min <- format(seq(start.min, 
                        length.out=24*60, by = "mins"),
                    '%H:%M')

## function to generate a trip sample
## min.dur : minimal duration of a trip
## max.dur : maximal duration of a trip
## min.trip : minimal number of trips that a user can do 

gen.Trip <- function(min.dur=3,max.dur=10,min.trip=100){
  ## gen number of trip
  n.trip <- sample(seq(min.trip,20),1)
  ## for each trip generate the durations
  durations <- rep(seq(1,n.trip),
                   times=sample(seq(min.dur,max.dur),
                                max(min.dur,n.trip),rep=TRUE))
  ## generate a vector of positions
  rr <- rle(durations)
  mm <- cumsum(rr$lengths)
  ## idrty part here
  pos <- sort(sample(seq(1,length(hours.min)-2*max(mm)),
              n.trip,rep=FALSE)) + mm
  ## assign each trip to each posistion  
  val <- vector(mode='integer',length(hours.min))
  for(x in seq_along(pos))
    val[seq(pos[x],length.out=rr$len[x])] <- rr$val[x]
  val
}

为 100 名司机生成行程

set.seed(1234)
nb.drivers <- 100
res <- replicate(nb.drivers,gen.Trip(),simplify=FALSE)
res <- do.call(rbind,res)
colnames(res) <- hours.min
rownames(res) <- paste0('driv',seq(nb.drivers))

宽幅

head(res[,10:30])
  ##       00:09 00:10 00:11 00:12 00:13 00:14 00:15 00:16 00:17 00:18 00:19
## driv1     0     0     0     0     0     0     1     1     1     1     1
## driv2     0     1     1     1     1     1     1     2     2     2     1
## driv3     0     0     0     0     0     0     0     0     0     0     0
## driv4     1     1     1     0     0     0     0     0     0     0     0
## driv5     0     0     0     0     0     0     0     0     0     0     1
## driv6     0     0     0     0     0     0     0     0     0     0     0
##       00:20 00:21 00:22 00:23 00:24 00:25 00:26 00:27 00:28 00:29
## driv1     1     1     0     0     2     2     2     2     2     2
## driv2     0     0     0     0     0     0     3     3     3     3
## driv3     0     0     0     0     0     0     0     0     0     0
## driv4     0     0     0     0     0     0     0     0     0     0
## driv5     1     1     1     1     1     1     1     1     0     0
## driv6     0     0     0     0     0     0     0     0     0     0

长格式

res.m <- melt(res)
head(res.m)
##    Var1  Var2 value
## 1 driv1 00:00     0
## 2 driv2 00:00     0
## 3 driv3 00:00     0
## 4 driv4 00:00     0
## 5 driv5 00:00     0
## 6 driv6 00:00     0

【讨论】:

  • 谢谢@agstudy。我将使用从真实驱动程序收集的数据(即我不会自己生成它),但此示例可用于在此处在更大的集合上验证我的模型。
  • @GeorgeK 使用我生成的数据回答您的问题有意义吗?
  • 对不起,我想我可能让大家对我的示例输出感到困惑。它与我提供的样本数据无关或结果。我只是想通过一些理论上的旅行(在最右边的列中解释)来展示给定一天的前 10 分钟的快照,它会是什么样子。
  • 嗨@agstudy。我已经尝试使用您提供的代码。我并不完全清楚(由于我的经验有限)您生成的数据(基于定义的规范等)如何转换为您的输出。我不确定它是否允许我使用收集的数据而不是生成的数据?
  • @GeorgeK 你是什么意思?生成的数据他们不复制收集的数据(链接的图片?)?如果没有,为什么?如果可以,我的问题很简单:收集的数据与您在此问题中提供的数据之间有什么关系? testdata 是否对收集的数据进行了转换?
【解决方案2】:

在此解决方案中,我读取了您的原始数据并将其格式化以获取我之前答案的生成数据。提供的数据仅限于司机的 22 次行程,但此处的重塑不受行程次数的限制。这个想法类似于用于生成样本数据的想法。我正在使用data.table,因为它可以方便地操作每个组的数据。

所以对于每一天,司机我都做以下事情:

  1. 创建一个长度为分钟数的零向量
  2. 使用 XXXstrip_start 和 XXXstrip_end 读取开始和结束位置。
  3. 创建序列 seq(start,end)
  4. 使用此序列将零更改为数字序列

这是我的代码:

start.min <- as.POSIXct("2013-01-01 00:00:00 PST")
hours.min <- format(seq(start.min, 
                        length.out=24*60, by = "mins"),
                    '%H:%M')
library(data.table)
diary <- read.csv("samplediary.csv",
                  stringsAsFactors=FALSE)
DT <- data.table(diary,key=c('id','veh_assigned','day'))

dat <- DT[, as.list({ .SD;nb.trip=sum_trips
           tripv <- vector(mode='integer',length(hours.min))
           if(sum_trips>0){
             starts = mget(paste0('X',seq(nb.trip),'_trip_start'))
             ends = mget(paste0('X',seq(nb.trip),'_trip_end'))
             ids <- mapply(function(x,y){
                                        seq(as.integer(x),as.integer(y))},
                           starts,ends,SIMPLIFY = FALSE)
             for (x in seq_along(ids))tripv[ids[[x]]] <- x
             }
            tripv
           }),
   by=c('id','day')]
setnames(x=dat,old=paste0('V',seq(hours.min)),hours.min)

如果您对前 10 个变量进行子集化,您会得到什么:

dat[1:10,1:10,with=FALSE]


       id day 00:00 00:01 00:02 00:03 00:04 00:05 00:06 00:07
 1: 3847339   1     0     0     0     0     0     0     0     0
 2: 3847384   1     0     0     0     0     0     0     0     0
 3: 3847436   1     0     0     0     0     0     0     0     0
 4: 3847439   1     0     0     0     0     0     0     0     0
 5: 3847510   1     0     0     0     0     0     0     0     0
 6: 3847536   1     0     0     0     0     0     0     0     0
 7: 3847614   1     0     0     0     0     0     0     0     0
 8: 3847683   1     0     0     0     0     0     0     0     0
 9: 3847841   1     0     0     0     0     0     0     0     0
10: 3847850   1     0     0     0     0     0     0     0     0

一个想法是创建数据的热图(至少每天),以获取一些直觉并查看重叠的驱动程序。这里有两种使用latticeggplot2 的方法,但首先我将使用reshape2 以长格式重塑数据

library(reshape2)
dat.m <- melt(dat,id.vars=c('id','day'))

然后我绘制热图以查看哪些驱动程序与其他驱动程序重叠:

library(lattice)
levelplot(value~as.numeric(variable)*factor(id),data=dat.m)

library(ggplot2)
ggplot(dat.m, aes(x=as.numeric(variable),y=factor(id)))+ 
        geom_tile(aes(fill = value)) +
  scale_fill_gradient(low="grey",high="blue")

【讨论】:

  • 哇!谢谢@agstudy。这正是我正在寻找的输出。你的解释也很清楚。非常感谢您分享您的时间和知识。我想我也会尝试用 R 来解决接下来的步骤。
  • @GeorgeK 欢迎您。我认为您也可以通过距离或/和速度更改tripv[ids[[x]]] &lt;- x(简单序列)以获得更多见解....
猜你喜欢
  • 2021-01-26
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-11-04
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多