【问题标题】:Calculate time difference and results are all 0计算时间差,结果全为0
【发布时间】:2020-02-26 22:58:52
【问题描述】:

我需要计算时间差并制作一个变量“持续时间”。示例数据如下所示:

                endTime                   startTime     user_id      categories duration
1      2019-02-22T19:17:02.618  019-02-22T19:16:58.377   10224   communication   0 secs
2      2019-02-23T12:19:01.055 2019-02-23T12:18:44.414   10224   communication   0 secs
3      2019-02-25T21:03:15.771 2019-02-25T21:03:06.961   10224 utility & tools   0 secs
4      2019-02-27T19:22:41.174 2019-02-27T19:22:32.246   10224   communication   0 secs

endTime 和 startTime 都使用 as.POSIXct 设置为日期格式。我在base R中使用difftime,代码如下:

dat$duration <- difftime(dat$startTime,dat$endTime)

所有持续时间的值为 0。我不明白为什么会这样。 我也检查了其他一些库(chron,lubridate)来计算这个,似乎它们只接受一个包含两个时间的字符串,而不是两个变量。将两个变量合并到一个字符串中对我来说似乎并不明智..有更简单的方法吗?谢谢!!

输入:

structure(list(battery = c(47L, 41L, 18L, 94L, 94L, 93L, 73L, 
73L, 47L, 49L), endTime = c("2019-02-22T19:17:02.618", "2019-02-23T12:19:01.055", 
"2019-02-25T21:03:15.771", "2019-02-27T19:22:41.174", "2019-02-27T19:22:53.256", 
"2019-02-27T23:51:16.407", "2019-03-02T20:18:28.090", "2019-03-02T20:18:43.488", 
"2019-03-19T13:07:16.993", "2019-03-19T12:16:36.962"), session = c(1550859371L, 
1550920714L, 1551124876L, 1551291720L, 1551291720L, 1551307871L, 
1551554295L, 1551554295L, 1552997232L, 1552994133L), startTime = c("2019-02-22T19:16:58.377", 
"2019-02-23T12:18:44.414", "2019-02-25T21:03:06.961", "2019-02-27T19:22:32.246", 
"2019-02-27T19:22:45.404", "2019-02-27T23:51:15.270", "2019-03-02T20:18:21.362", 
"2019-03-02T20:18:37.066", "2019-03-19T13:07:15.348", "2019-03-19T12:15:38.440"
), user_id = c(10224L, 10224L, 10224L, 10224L, 10224L, 10224L, 
10224L, 10224L, 10224L, 10224L), categories = structure(c(1L, 
1L, 6L, 1L, 2L, 2L, 6L, 1L, 2L, 1L), .Label = c("communication", 
"games & entertainment", "lifestyle", "news & information outlet", 
"social network", "utility & tools"), class = "factor"), duration = structure(c(0, 
0, 0, 0, 0, 0, 0, 0, 0, 0), class = "difftime", units = "secs")), row.names = c(NA, 
10L), class = "data.frame")

【问题讨论】:

  • 您确定变量是Date 格式吗?如果是,difftime 很可能会返回 days 中的差异,在您的示例中等于零。
  • 如果您包含一个简单的reproducible example,其中包含可用于测试和验证可能解决方案的示例输入和所需输出,则更容易为您提供帮助。分享您的数据的dput(),以便我们可以准确查看其中的内容。仅仅看到数据如何打印到屏幕上是不够的。
  • @JDG 嗨,日期对我的分析并不那么重要,这就是为什么我试图删除它并只留下小时/分钟/秒,所以它不会计算日差。但是,它不起作用...您能否推荐如何将时间设置为日期格式?这是我第一次使用时间数据...谢谢!
  • 您刚刚以文本格式粘贴了它。我们对此无能为力。粘贴整个dput() 输出,即它看起来像structure(...)
  • 您的日期列很可能是字符串或因子,而不是日期/时间对象。如果为 True,则需要使用 as.POSIXct() 函数进行转换,然后使用 difftime()。这是一个很好的例子:stackoverflow.com/questions/21667212/…

标签: r dataframe posixct


【解决方案1】:

以下内容对我有用。 difftime 不为零。

op_digs <- options(digits.secs = 3)

dat$endTime <- as.POSIXct(dat$endTime, format = "%Y-%m-%dT%H:%M:%OS")
dat$startTime <- as.POSIXct(dat$startTime, format = "%Y-%m-%dT%H:%M:%OS")
difftime(dat$startTime,dat$endTime)
#Time differences in secs
#[1]  -4.241 -16.641  -8.810  -8.928

dat$duration <- difftime(dat$startTime,dat$endTime)

options(digits.secs = op_digs)

数据。

dat <- read.table(text = "
                endTime                   startTime     user_id      categories duration
1      2019-02-22T19:17:02.618 2019-02-22T19:16:58.377   10224   communication   '0 secs'
2      2019-02-23T12:19:01.055 2019-02-23T12:18:44.414   10224   communication   '0 secs'
3      2019-02-25T21:03:15.771 2019-02-25T21:03:06.961   10224 'utility & tools'   '0 secs'
4      2019-02-27T19:22:41.174 2019-02-27T19:22:32.246   10224   communication   '0 secs'
", header = TRUE)

【讨论】:

  • 谢谢!当我运行完全相同的代码时,当我打开检查数据集时它给了我一个错误,说'r error 4(R code execution error)'。当我删除最后一行'options(digits.secs = op_digs)'时,它起作用了。但是如果我也删除第一行'op_digs
【解决方案2】:

这里有一个基于dplyr的解决方案,工作流更流畅:

library(dplyr)
dat %>%
  mutate_at(vars(startTime, endTime), ~as.POSIXct(strptime(.x, format = c("%Y-%m-%dT%H:%M:%OS")))) %>%
  mutate(duration = endTime - startTime)

   battery             endTime    session           startTime user_id            categories    duration
1       47 2019-02-22 19:17:02 1550859371 2019-02-22 19:16:58   10224         communication  4.241 secs
2       41 2019-02-23 12:19:01 1550920714 2019-02-23 12:18:44   10224         communication 16.641 secs
3       18 2019-02-25 21:03:15 1551124876 2019-02-25 21:03:06   10224       utility & tools  8.810 secs
4       94 2019-02-27 19:22:41 1551291720 2019-02-27 19:22:32   10224         communication  8.928 secs
5       94 2019-02-27 19:22:53 1551291720 2019-02-27 19:22:45   10224 games & entertainment  7.852 secs
6       93 2019-02-27 23:51:16 1551307871 2019-02-27 23:51:15   10224 games & entertainment  1.137 secs
7       73 2019-03-02 20:18:28 1551554295 2019-03-02 20:18:21   10224       utility & tools  6.728 secs
8       73 2019-03-02 20:18:43 1551554295 2019-03-02 20:18:37   10224         communication  6.422 secs
9       47 2019-03-19 13:07:16 1552997232 2019-03-19 13:07:15   10224 games & entertainment  1.645 secs
10      49 2019-03-19 12:16:36 1552994133 2019-03-19 12:15:38   10224         communication 58.522 secs

基本上,您首先将列 endTime 和 startTime 转换为适当的格式 (POSIXct),然后使用简单的减法。

【讨论】:

    【解决方案3】:

    以下是使用 lubridate 和 dplyr 包解决此问题的方法。它将提供与上述相同的结果。只是方式不同

    df<-structure(list(battery = c(47L, 41L, 18L, 94L, 94L, 93L, 73L, 73L, 47L, 49L), 
                       endTime = c("2019-02-22T19:17:02.618", "2019-02-23T12:19:01.055", "2019-02-25T21:03:15.771", "2019-02-27T19:22:41.174", "2019-02-27T19:22:53.256", "2019-02-27T23:51:16.407", "2019-03-02T20:18:28.090", "2019-03-02T20:18:43.488", "2019-03-19T13:07:16.993", "2019-03-19T12:16:36.962"), 
                       session = c(1550859371L,1550920714L, 1551124876L, 1551291720L, 1551291720L, 1551307871L,1551554295L, 1551554295L, 1552997232L, 1552994133L),
                       startTime = c("2019-02-22T19:16:58.377",  "2019-02-23T12:18:44.414", "2019-02-25T21:03:06.961", "2019-02-27T19:22:32.246", "2019-02-27T19:22:45.404", "2019-02-27T23:51:15.270", "2019-03-02T20:18:21.362", "2019-03-02T20:18:37.066", "2019-03-19T13:07:15.348", "2019-03-19T12:15:38.440"),
                       user_id = c(10224L, 10224L, 10224L, 10224L, 10224L, 10224L, 10224L, 10224L, 10224L, 10224L), 
                       categories = structure(c(1L, 1L, 6L, 1L, 2L, 2L, 6L, 1L, 2L, 1L), .Label = c("communication", "games & entertainment", "lifestyle", "news & information outlet", "social network", "utility & tools"), class = "factor"), 
                       duration = structure(c(0, 0, 0, 0, 0, 0, 0, 0, 0, 0), class = "difftime", units = "secs")), row.names = c(NA, 10L), class = "data.frame")
    
    library(dplyr)
    library(lubridate)
    df<-df %>%
      mutate(endTime = ymd_hms(endTime), startTime = ymd_hms(startTime)) %>%
      mutate(duration = endTime - startTime)
    df2
    

    这是输出

       battery             endTime    session           startTime user_id            categories    duration
    1       47 2019-02-22 19:17:02 1550859371 2019-02-22 19:16:58   10224         communication  4.241 secs
    2       41 2019-02-23 12:19:01 1550920714 2019-02-23 12:18:44   10224         communication 16.641 secs
    3       18 2019-02-25 21:03:15 1551124876 2019-02-25 21:03:06   10224       utility & tools  8.810 secs
    4       94 2019-02-27 19:22:41 1551291720 2019-02-27 19:22:32   10224         communication  8.928 secs
    5       94 2019-02-27 19:22:53 1551291720 2019-02-27 19:22:45   10224 games & entertainment  7.852 secs
    6       93 2019-02-27 23:51:16 1551307871 2019-02-27 23:51:15   10224 games & entertainment  1.137 secs
    7       73 2019-03-02 20:18:28 1551554295 2019-03-02 20:18:21   10224       utility & tools  6.728 secs
    8       73 2019-03-02 20:18:43 1551554295 2019-03-02 20:18:37   10224         communication  6.422 secs
    9       47 2019-03-19 13:07:16 1552997232 2019-03-19 13:07:15   10224 games & entertainment  1.645 secs
    10      49 2019-03-19 12:16:36 1552994133 2019-03-19 12:15:38   10224         communication 58.522 secs
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多