【问题标题】:Dataframe Dealing With Missing Time Series Data处理缺失时间序列数据的数据框
【发布时间】:2016-11-04 16:28:04
【问题描述】:

我的 DataFrame 是一个数组时间序列,在大约 60 天内每分钟拍摄一次。

  1. 首先我想将 df 分割成 24 小时周期。

  2. 然后我想将某些属性绘制为瀑布图,折线图相互叠加。

我正在考虑在 for 循环中使用 iloc 来执行此操作,因为 df 行是按时间索引的,这意味着每天有 3600 行。问题是我不知道如何将每个分配给一个变量。

for i in range(58)
     df = timethingdf.iloc[809+i*3600:809+(i+1)*3600]

如您所见,我希望 df 对于我用它制作的 58 个 dfs 中的每一个都不同。

而且我不知道如何制作图表。

【问题讨论】:

    标签: python pandas matplotlib


    【解决方案1】:

    我认为你应该是这个意思:

    for i in range(58)
        df = timethingdf.iloc[809+i*3600:809+(i+1)*3600]
        # Doing something with `df`
    

    【讨论】:

      【解决方案2】:

      我想你想要的是TimeGrouper:

      data = {'date':['2004-1-2:10:10:00', '2004-1-2:10:11:00', '2004-1-1:11:11:00', '2004-1-1:11:13:00'], 'foo':[5,6,7,8]}
      df = pd.DataFrame(data)
      df['date'] = pd.to_datetime(df['date'], format='%Y-%m-%d:%H:%M:%S')
      df = df.set_index('date')
      grouped = df.groupby(pd.TimeGrouper('24H')).sum()
      
      In [7]: grouped
      Out[8]:
                  foo
      date
      2004-01-01   15
      2004-01-02   11
      

      然后,您可以将 .sum() 替换为您想在分组子集上使用的任何聚合器。

      【讨论】:

        猜你喜欢
        • 2017-09-26
        • 2015-12-18
        • 1970-01-01
        • 1970-01-01
        • 2022-01-06
        • 2016-11-21
        • 1970-01-01
        • 1970-01-01
        • 2021-05-27
        相关资源
        最近更新 更多