【问题标题】:How to fill missing timestamps for Time column for a date in pandas如何为熊猫中的日期填充时间列的缺失时间戳
【发布时间】:2018-08-17 16:00:06
【问题描述】:

我有一个时间序列数据如下:

print(df)

    ric     datel       timel        val
0   xyz     2017-01-01  09:00:00     2
1   xyz     2017-01-01  09:04:00     5
2   xyz     2017-01-01  09:37:00     6

现在我必须将缺失的时间戳填充到09:45:00。

预期输出:

    ric     datel       timel        val
0   xyz     2017-01-01  09:00:00     2
1   xyz     2017-01-01  09:01:00     nan
2   xyz     2017-01-01  09:02:00     nan
3   xyz     2017-01-01  09:03:00     nan
4   xyz     2017-01-01  09:04:00     5
...
...
37  xyz     2017-01-01  09:37:00      6
...
...
45  xyz     2017-01-01  09:45:00      nan

我尝试了什么:

df1=df.resample("1 min", on ='datel').first()

输出如下:

              ric   datel       timel     val
datel                   
2017-01-01  xyz     2017-01-01  09:00:00    2

还尝试使用pd.date_range,但它主要适用于日期时间列。 我有两个不同的日期和时间列。有没有办法在不将日期和列组合成日期时间的情况下实现这一点?

【问题讨论】:

    标签: python pandas time-series


    【解决方案1】:

    主要思想是使用reindex by times 创建 date_range:

    df['timel'] = pd.to_datetime(df['timel']).dt.time
    start = pd.to_datetime(str(df['timel'].min()))
    end = pd.to_datetime('09:45:00')
    dates = pd.date_range(start=start, end=end, freq='1Min').time
    #print (dates)
    
    df = df.set_index('timel').reindex(dates).reset_index().reindex(columns=df.columns)
    cols = df.columns.difference(['val'])
    df[cols] = df[cols].ffill()
    print (df.head())
       ric       datel     timel  val
    0  xyz  2017-01-01  09:00:00  2.0
    1  xyz  2017-01-01  09:01:00  NaN
    2  xyz  2017-01-01  09:02:00  NaN
    3  xyz  2017-01-01  09:03:00  NaN
    4  xyz  2017-01-01  09:04:00  5.0
    

    与resample类似的解决方案:

    df['timel'] = pd.to_datetime(df['timel'])
    
    #if missing row with 09:45:00 add it
    if not (df['timel']  == pd.to_datetime('09:45:00')).any():
        df.loc[len(df.index), 'timel'] = pd.to_datetime('09:45:00')
    
    df=df.set_index('timel').resample("1min").first().reset_index().reindex(columns=df.columns)
    cols = df.columns.difference(['val'])
    df[cols] = df[cols].ffill()
    df['timel'] = df['timel'].dt.time
    print (df.head())
       ric       datel     timel  val
    0  xyz  2017-01-01  09:00:00  2.0
    1  xyz  2017-01-01  09:01:00  NaN
    2  xyz  2017-01-01  09:02:00  NaN
    3  xyz  2017-01-01  09:03:00  NaN
    4  xyz  2017-01-01  09:04:00  5.0
    

    【讨论】:

    • @AkshayNevrekar - 很高兴能帮上忙!
    • 这正是我要找的……非常感谢!
    【解决方案2】:

    使用 date_range 生成日期后,您可以使用类似于下面的函数对其进行拆分。

    返回值可以输入到df中

    从日期时间导入日期时间

    def split_datetime(date_with_time):
        """
        This function will return date and time from datetime input
        """
        date_with_time = date_with_time.split(' ')
        date = date_with_time[0]
        time = date_with_time[1].split('.')[0]
        return date, time
    
    #Eg:                   
    date, time = split_datetime(str(datetime.now()))
    

    【讨论】:

      猜你喜欢
      • 2018-04-24
      • 2020-06-18
      • 1970-01-01
      • 2018-11-10
      • 2017-02-23
      • 2020-09-13
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多