【问题标题】:Consecutive/Unbroken occurrences in time series时间序列中连续/不间断的事件
【发布时间】:2018-12-18 16:39:24
【问题描述】:

我正在使用 R 中的数据表,其中包含有关美国杂货店所售产品的季度信息。特别是,有一个日期列、一个商店列和一个产品列。例如,这里是数据的一个(非常小的)子集:

Date           StoreID       ProductID
2000-03-31     10001         20001       
2000-03-31     10001         20002
2000-03-31     10002         20001
2000-06-30     10001         20001

对于每家商店中的每种产品,我想知道该产品在该日期之前在该商店中销售了多少个连续个季度。例如,如果我们只查看在特定商店出售的订书机,我们会:

Date           StoreID       ProductID
2000-03-31     10001         20001       
2000-06-30     10001         20001
2000-09-30     10001         20001
2000-12-31     10001         20001      
2001-06-30     10001         20001
2001-09-30     10001         20001
2001-12-31     10001         20001

假设这是 StoreID 和 ProductID 组合的所有数据,我想分配一个新变量:

Date           StoreID       ProductID     V
2000-03-31     10001         20001         1
2000-06-30     10001         20001         2
2000-09-30     10001         20001         3
2000-12-31     10001         20001         4
2001-06-30     10001         20001         1
2001-09-30     10001         20001         2
2001-12-31     10001         20001         3
2002-03-31     10001         20001         4
2002-06-30     10001         20001         5
2002-09-30     10001         20001         6
2002-12-31     10001         20001         7
2004-03-30     10001         20001         1
2004-06-31     10001         20001         2

请注意,我们在 2000 年第四季度之后结转,因为该产品在 2001 年第一季度没有售出。此外,我们在 2002 年第四季度之后结转,因为该产品在 2003 年第一季度没有售出。下一次售出该产品是 2004 年第一季度,它被分配了一个 1。

我遇到的问题是,我的实际数据集非常大(大约 1000 万行),所以这需要有效地完成。我能想出的唯一技术效率极低。任何建议将不胜感激。

【问题讨论】:

  • 为什么 V 在第 5 行重置为 1?是因为时差太大还是因为过年?另外,请使用dput 函数添加更多数据(多个商店和产品)。
  • 它重置,因为产品 20001 在商店 10001 一个季度内没有售出。请注意日期跳过。
  • 我的解决方案是否有效或有什么我可以改进的地方?
  • @PoGibas 我现在正在评估它。看起来很有前途!

标签: r datetime


【解决方案1】:

您可以使用自定义函数计算季度之间的差异。

# Load data.table
library(data.table)
# Set data as a data.table object
setDT(data)
# Set key as it might be big data
setkey(data, StoreID, ProductID)

consecutiveQuarters <- function(date, timeGap = 14) {
    # Calculate difference in dates 
    # And check if this difference is less than 14 weeks
    shifts <- cumsum(c(FALSE, abs(difftime(date[-length(date)], date[-1], units = "weeks")) > timeGap))
    # Generate vector from 1 to number of consecutive quarters
    ave(shifts, shifts, FUN = seq_along)
}

# Calculate consecutive months my storeID and productID
data[, V := consecutiveQuarters(Date), .(StoreID, ProductID)]

【讨论】:

  • 所以这不太奏效。对于给定的商店,假设连续 10 个季度销售相同的产品(为简单起见,从给定年份的第一季度开始)。查看日期,而不是得到 1,2,3,4,5,6,...,10,这个代码给了我 1,2,3,4,1,2,3,4,1,2。发生这种情况是因为年份翻转时的季度差为 -1。如果我们通过取模 4 的差值来更改 ContinuousQuarters 函数中“移位”的定义,这会解决问题,但会增加一个新问题。该代码会将 Q1Y1 和 Q2Y100 视为连续的。我们必须以某种方式合并年份吗?想法?
  • @Joe 您能否在您的问题中添加这样的示例(即必须返回 1:10 的连续年份)。您可以为此使用函数dput()
  • 我添加了这样一个例子(但只有 1-7 而不是 1-10)。我还添加了一个示例,说明引入模运算会导致的新问题。
  • @Joe 我编辑了我的答案。我将其更改为使用函数difftime。我们用它计算周差(最大允许单位)并检查这个差是否
  • 效果很好。非常感谢您的帮助!
【解决方案2】:

我从您的问题中了解到,您确实需要将您的 V 列作为年度季度,而不是每个季度的产品总和。你可以使用类似的东西。

# to_quarters returns year's quarter of given date in character string
# base on reg exp    
to_quarters <- function(date_string) {
  month <- as.numeric(substr(date_string, 6, 7))
  as.integer((month - 1) / 3) + 1
}

# with tidyverse library
library(tidyverse)
# your data as tibble format of data frame
data_set_tibble <- as.tibble(YOUR_DATA)
# here you create your table 
data_set_tibble %>% mutate(V = to_quarters(Date) %>% as.integer())


# alterative with data.table library
library(data.table)
# your data as data.table format of data frame
data_set  <- as.data.table(YOUR_DATA)
# here you create your table 
data_set[,.(Date, StoreID, ProductID, V = to_quarters(Date))]

对于 tidyverse 和 data.table 的性能是相同的,在我的例子中,5 000 000 行在 12 秒内工作

【讨论】:

  • 这不是我需要的。我想要产品在商店中销售的连续季度数。不是一年中的那个季度。
【解决方案3】:

创建一个变量,如果产品在一个季度内售出,则为 1,否则为 0。对变量进行排序,使其从当前开始并及时向后移动。

将此类变量的累积和与相同长度的序列进行比较。当销售额降至零时,累积总和将不再等于序列。将累积和等于序列的次数相加,这将说明连续季度销售额为正数。

data <- data.frame(
  quarter = c(1, 2, 3, 4, 1, 2, 3, 4),
  store = as.factor(c(1, 1, 1, 1, 1, 1, 1, 1)),
  product = as.factor(c(1, 1, 1, 1, 2, 2, 2, 2)),
  numsold = c(5, 6, 0, 1, 7, 3, 2, 14)
)


sortedData <- data[order(-data$quarter),]

storeValues <- c("1")
productValues <- c("1","2")

dataConsec <- data.frame(store = NULL, product = NULL, ConsecutiveSales = NULL)

for (storeValue in storeValues ){
  for(productValue in productValues){

    prodSoldinQuarter <- 
      as.numeric(sortedData[sortedData$store == storeValue &
                        sortedData$product == productValue,]$numsold > 0)

    dataConsec <- rbind(dataConsec,
                        data.frame(
                          store = storeValue,
                          product = productValue,
                          ConsecutiveSales = 
                            sum(as.numeric(cumsum(prodSoldinQuarter) == 
                                     seq(1,length(prodSoldinQuarter)) 
                                    ))
                          ))

  }
} 

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2018-08-27
    • 2014-03-09
    • 2023-03-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-07-02
    相关资源
    最近更新 更多