【发布时间】:2019-08-14 13:54:02
【问题描述】:
我正在尝试获取一年中每一天的库存汽车数量,以及该日期每辆车的库存天数。
我有完整的移动历史记录(每辆车进出库存的时间戳 - 出租、出售、维修等)。像这样:
car in out status_id operation
PZR4010 08/02/2018 08:55 08/02/2018 16:29 12 out_stock
QRX0502 07/02/2018 09:00 07/02/2018 10:28 7 in_stock
PYR8269 06/02/2018 17:10 09/02/2018 21:22 12 in_stock
QRG6455 06/02/2018 12:39 8 sold
QRU1867 08/02/2018 08:00 09/02/2018 11:07 12 in_stock
PZR8528 06/02/2018 17:51 07/02/2018 07:46 10 out_stock
PZR7184 06/02/2018 16:00 08/02/2018 12:10 7 in_stock
PZR0386 08/02/2018 09:02 14/02/2018 14:53 10 out_stock
PZR8600 06/02/2018 16:00 07/02/2018 07:34 7 in_stock
PZR1787 06/02/2018 17:02 20/02/2018 17:33 12 in_stock
因此,对于每辆车,我必须加入它的整个连续库存时间,以了解它处于该状态的时间。
例如:
car in out status_id operation
QRX0502 08/02/2018 08:55 09/02/2018 16:29 7 in_stock
QRX0502 07/02/2018 09:00 08/02/2018 08:55 7 in_stock
QRX0502 06/02/2018 17:10 07/02/2018 09:00 7 in_stock
将变得简单:
car in out status_id operation
QRX0502 06/02/2018 17:10 09/02/2018 16:29 7 in_stock
捕获“in”列中的最小时间戳和“out”列中的最大时间戳。
我尝试过使用 groupby + shift:
#'mov' is the dataframe with all the stock movements
# I create a columns to better filter on the groupby
mov['aux']=mov['car']+" - "+mov['operation']
#creating the base dataframe to be the output
hist_mov=pd.DataFrame(columns=list(mov.columns))
for line, operation in mov.groupby(mov['aux'].ne(mov['aux'].shift()).cumsum()):
g_temp=operation.groupby(['car','operation',
'aux']).agg({'in':'min','out':'max'}).reset_index()
hist_mov=hist_mov.append(g_temp,sort=True)
问题是整个数据库运行大约需要 16 个小时,而且我必须每天运行它来更新库存状态。
我想构建类似的东西:
添加到历史记录的每一行都会检查它是否与我的新基础 (hist_mov) 中的任何一行连续。如果是这样,请更新该行。如果没有,请添加为新行。
有什么想法吗?谢谢!
【问题讨论】:
标签: python pandas pandas-groupby