【问题标题】:How to fill missing date in timeSeries如何在时间序列中填写缺失的日期
【发布时间】:2018-05-21 08:01:11
【问题描述】:

我的数据如下所示:

有每日记录,除了从 2017-06-12 到 2017-06-16 的间隔。

df2['timestamp'] = pd.to_datetime(df['timestamp'])
df2['timestamp'] = df2['timestamp'].map(lambda x: 
datetime.datetime.strftime(x,'%Y-%m-%d'))
df2 = df2.convert_objects(convert_numeric = True)
df2 = df2.groupby('timestamp', as_index = False).sum()

我需要用所有字段的值来填补这个缺失的空白和其他空白(例如timestamp、temperature、humidity、light、pressure、speed、battery_voltage 等。 .).

如何使用 Pandas 完成此任务?

这是我以前做过的

weektime = pd.date_range(start = '06/04/2017', end = '12/05/2017', freq = 'W-SUN')
df['week'] = 'nan'
df['weektemp'] = 'nan'
df['weekhumidity'] = 'nan'
df['weeklight'] = 'nan'
df['weekpressure'] = 'nan'
df['weekspeed'] = 'nan'
df['weekbattery_voltage'] = 'nan'

for i in range(0,len(weektime)):
    df['week'][i+1] = weektime[i]
    df['weektemp'][i+1] = df['temperature'].iloc[7*i+1:7*i+7].sum()
    df['weekhumidity'][i+1] = df['humidity'].iloc[7*i+1:7*i+7].sum()
    df['weeklight'][i+1] = df['light'].iloc[7*i+1:7*i+7].sum()
    df['weekpressure'][i+1] = df['pressure'].iloc[7*i+1:7*i+7].sum()
    df['weekspeed'][i+1] = df['speed'].iloc[7*i+1:7*i+7].sum()
    df['weekbattery_voltage'][i+1] = 
df['battery_voltage'].iloc[7*i+1:7*i+7].sum()
     i = i + 1

sum 的值不正确。因为 2017-06-17 的值是 2017-06-12 到 2017-06-16 的总和。我不想再次添加它们。这个差距不仅是一个时期的差距。我想把它们都填满。

【问题讨论】:

  • 有很多关于填补缺失日期时间的链接。 Here is one 但还有更多
  • 您没有解释,填充缺失数据的基本假设是什么。除时间外的所有参数都是非线性的,并且在五个缺失数据点的时间过程中发生显着变化。
  • 我需要将每个第一个数据到第 7 个数据的 7 天数据相加。(每周)(例如第一个数据集是:df.iloc[0:6].sum[] 和第二个数据集是 df.iloc[7:13].sum() 等等)。如果日期时间有间隔,则sum的值会出错,其他的也会出错。
  • 你的数据点无论如何都是错误的,因为你的数据点包含纯猜测的 6/7。与其假装发明的数据点反映了现实,不如说清楚数据点丢失了。例外情况是您可以证明如何预测缺失值。但是例如查看pressure 数据点。他们会在缺失的一周内以线性方式上升吗?他们会在缺失的一周开始或结束时突然上升吗?你不知道。

标签: python pandas


【解决方案1】:

这是我写的一个函数,可能对你有帮助。它会在时间上寻找不一致的跳跃并填充它们。使用此功能后,尝试使用线性插值函数(pandas 有一个很好的)来填充您的空数据值。注意:Numpy 数组的迭代和操作比 Pandas 数据帧快得多,这就是我在两者之间切换的原因。

import numpy as np
import pandas as pd

data_arr = np.array(your_df)
periodicity = 'daily'

def fill_gaps(data_arr, periodicity):
    rows = data_arr.shape[0]
    data_no_gaps = np.copy(data_arr) #avoid altering the thing you're iterating over
    data_no_gaps_idx = 0

    for row_idx in np.arange(1, rows): #iterate once for each row (except the first record; nothing to compare)
        oldtimestamp_str = str(data_arr[row_idx-1, 0]) 
        oldtimestamp = np.datetime64(oldtimestamp_str)  

        currenttimestamp_str = str(data_arr[row_idx, 0])
        currenttimestamp = np.datetime64(currenttimestamp_str)

        period = currenttimestamp - oldtimestamp

        if period != np.timedelta64(900,'s') and period != np.timedelta64(3600,'s') and period != np.timedelta64(86400,'s'):                                
            if periodicity == 'quarterly':
                desired_period = 900
            elif periodicity == 'hourly':
                desired_period = 3600
            elif periodicity == 'daily':
                desired_period = 86400

            periods_missing = int(period / np.timedelta64(desired_period,'s'))
            for missing in np.arange(1, periods_missing):
                new_time_orig = str(oldtimestamp + missing*(np.timedelta64(desired_period,'s')))
                new_time = new_time_orig.replace('T', ' ')
                data_no_gaps = np.insert(data_no_gaps, (data_no_gaps_idx + missing), 
                                 np.array((new_time, np.nan, np.nan, np.nan, np.nan, np.nan)), 0) # INSERT VALUES YOU WANT IN THE NEW ROW

            data_no_gaps_idx += (periods_missing-1) #incriment the index (zero-based => -1) in accordance with added rows

        data_no_gaps_idx += 1 #allow index to change as we iterate over original data array (main for loop)

    #create a dataframe:
    data_arr_no_gaps = pd.DataFrame(data=data_no_gaps, index=None,columns=['Time', 'temp', 'humidity', 'light', 'pressure', 'speed'])

    return data_arr_no_gaps  

【讨论】:

    【解决方案2】:

    填补时间空白和空值

    使用下面的函数确保预期的日期序列存在,然后使用前向填充填充空值。

    import pandas as pd
    import os
    
    def fill_gaps_and_nulls(df, freq='1D'):
        '''
        General steps:
            A) check for extra dates (out of expected frequency/sequence)
            B) check for missing dates (based on expected frequency/sequence)
            C) use forwardfill to fill nulls
            D) use backwardfill to fill remaining nulls
            E) append to file    
        '''
        
        #rename the timestamp to 'date'
        df.rename(columns={"timestamp": "date"})
        
        #sort to make indexing faster
        df = df.sort_values(by=['date'], inplace=False)
        
        #create an artificial index of dates at frequency = freq, with the same beginning and ending as the original data
        all_dates = pd.date_range(start=df.date.min(), end=df.date.max(), freq=freq)
    
        #record column names
        df_cols = df.columns
    
        #delete ffill_df.csv so we can begin anew
        try:
            os.remove('ffill_df.csv')
        except FileNotFoundError:
            pass
    
        #check for extra dates and/or dates out of order. print warning statement for log
        extra_dates = set(df.date).difference(all_dates)
        
        #if there are extra dates (outside of expected sequence/frequency), deal with them
        if len(extra_dates) > 0:
            #############################
            #INSERT DESIRED BEHAVIOR HERE        
            print('WARNING: Extra date(s):\n\t{}\n\t    Shifting highlighted date(s) back by 1 day'.format(extra_dates))
            for date in extra_dates:
                #shift extra dates back one day
                df.date[df.date == date] = date - pd.Timedelta(days=1)
            #############################
    
        #check the artificial date index against df to identify missing gaps in time and fill them with nulls
        gaps = all_dates.difference(set(df.date))
        print('\n-------\nWARNING: Missing dates: {}\n-------\n'.format(gaps))
    
        #if there are time gaps, deal with them
        if len(gaps) > 0:
            #initialize df of correct size, filled with nulls
            gaps_df = pd.DataFrame(index=gaps, columns=df_cols.drop('date')) #len(index) sets number of rows
            #give index a name
            gaps_df.index.name = 'date'
            #add the region and type
            gaps_df.region = r
            gaps_df.type = t
            #remove that index so gaps_df and df are compatible
            gaps_df.reset_index(inplace=True)                
            #append gaps_df to df
            new_df = pd.concat([df, gaps_df])
            #sort on date
            new_df.sort_values(by='date', inplace=True)
    
        #fill nulls
        new_df.fillna(method='ffill', inplace=True)
        new_df.fillna(method='bfill', inplace=True)
    
        #append to file
        new_df.to_csv('ffill_df.csv', mode='a', header=False, index=False) 
        
        return df_cols, regions, types, all_dates
    

    【讨论】:

      猜你喜欢
      • 2011-04-03
      • 1970-01-01
      • 2018-04-24
      • 2021-07-18
      • 1970-01-01
      • 2017-05-06
      • 1970-01-01
      • 2014-10-27
      • 2018-06-24
      相关资源
      最近更新 更多