【发布时间】:2014-09-09 01:03:00
【问题描述】:
我正在尝试编写一个函数,该函数采用连续的时间序列并返回一个数据结构,该数据结构描述了数据中任何缺失的间隙(例如,带有“start”和“end”列的 DF)。对于时间序列来说,这似乎是一个相当普遍的问题,但尽管搞乱了 groupby、diff 等 - 并探索了 SO - 我无法想出比下面更好的方法。
使用矢量化操作来保持效率是我的首要任务。必须有一个使用矢量化操作的更明显的解决方案——不是吗?感谢您的帮助,伙计们。
import pandas as pd
def get_gaps(series):
"""
@param series: a continuous time series of data with the index's freq set
@return: a series where the index is the start of gaps, and the values are
the ends
"""
missing = series.isnull()
different_from_last = missing.diff()
# any row not missing while the last was is a gap end
gap_ends = series[~missing & different_from_last].index
# count the start as different from the last
different_from_last[0] = True
# any row missing while the last wasn't is a gap start
gap_starts = series[missing & different_from_last].index
# check and remedy if series ends with missing data
if len(gap_starts) > len(gap_ends):
gap_ends = gap_ends.append(series.index[-1:] + series.index.freq)
return pd.Series(index=gap_starts, data=gap_ends)
为了记录,Pandas==0.13.1,Numpy==1.8.1,Python 2.7
【问题讨论】:
标签: python numpy pandas vectorization