【问题标题】:Describing gaps in a time series pandas描述时间序列 pandas 中的空白
【发布时间】:2014-09-09 01:03:00
【问题描述】:

我正在尝试编写一个函数,该函数采用连续的时间序列并返回一个数据结构,该数据结构描述了数据中任何缺失的间隙(例如,带有“start”和“end”列的 DF)。对于时间序列来说,这似乎是一个相当普遍的问题,但尽管搞乱了 groupby、diff 等 - 并探索了 SO - 我无法想出比下面更好的方法。

使用矢量化操作来保持效率是我的首要任务。必须有一个使用矢量化操作的更明显的解决方案——不是吗?感谢您的帮助,伙计们。

import pandas as pd


def get_gaps(series):
    """
    @param series: a continuous time series of data with the index's freq set
    @return: a series where the index is the start of gaps, and the values are
         the ends
    """
    missing = series.isnull()
    different_from_last = missing.diff()

    # any row not missing while the last was is a gap end        
    gap_ends = series[~missing & different_from_last].index

    # count the start as different from the last
    different_from_last[0] = True

    # any row missing while the last wasn't is a gap start
    gap_starts = series[missing & different_from_last].index        

    # check and remedy if series ends with missing data
    if len(gap_starts) > len(gap_ends):
         gap_ends = gap_ends.append(series.index[-1:] + series.index.freq)

    return pd.Series(index=gap_starts, data=gap_ends)

为了记录,Pandas==0.13.1,Numpy==1.8.1,Python 2.7

【问题讨论】:

    标签: python numpy pandas vectorization


    【解决方案1】:

    这个问题可以转化为查找列表中的连续数字。找出series为null的所有索引,如果(3,4,5,6)的run全部为null,则只需要提取开始和结束(3,6)

    import numpy as np
    import pandas as pd
    from operator import itemgetter
    from itertools import groupby
    
    
    # create an example 
    data = [2, 3, 4, 5, 12, 13, 14, 15, 16, 17]
    s = pd.series( data, index=data)
    s = s.reindex(xrange(18))
    print find_gap(s)  
    
    
    def find_gap(s): 
        """ just treat it as a list
        """ 
        nullindex = np.where( s.isnull())[0]
        ranges = []
        for k, g in groupby(enumerate(nullindex), lambda (i,x):i-x):
            group = map(itemgetter(1), g)
            ranges.append((group[0], group[-1]))
        startgap, endgap = zip(* ranges) 
        return pd.series( endgap, index= startgap )
    

    参考:Identify groups of continuous numbers in a list

    【讨论】:

    • 谢谢——你说得对,有很棒的 itertools 等解决方案。尽管我真的希望将其保持在 numpy 级别以提供尽可能多的优化,但我应该更清楚。另外,我不确定这是否真的是比上述更清洁的解决方案。
    • 也许你想看看这个解决方案:stackoverflow.com/questions/7352684/… unutbu 的那个
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2021-11-12
    • 1970-01-01
    • 1970-01-01
    • 2019-02-05
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多