【问题标题】:pandas time series average monthly volume熊猫时间序列平均每月交易量
【发布时间】:2018-10-15 20:41:51
【问题描述】:

我有每天一次的 csv 时间序列数据和累计销售。与此类似

01-01-2010 12:10:10      50.00
01-02-2010 12:10:10      80.00
01-03-2010 12:10:10      110.00
.
. for each dat of 2010
.
01-01-2011 12:10:10      2311.00
01-02-2011 12:10:10      2345.00
01-03-2011 12:10:10      2445.00
.
. for each dat of 2011
.

and so on.  

我希望获得每年每个月的月销售额(最大值 - 最小值)。因此,在过去的 5 年中,我将有 5 个 1 月值(最大值 - 最小值)、5 个 2 月值(最大值 - 最小值)......等等

一旦我有了这些,接下来我会得到 1 月的(5 年平均),2 月的 5 年平均......等等。

现在,我通过对原始 df [年/月] 进行切片,然后对一年中的特定月份进行平均。

我希望使用时间序列 resample() 方法,但我目前坚持告诉 PD 在 [从今天起的过去 10 年] 中每个月每月(最大 - 最小)采样。然后链入一个 .mean()

任何关于使用 resample() 的有效方法的建议都将不胜感激。

【问题讨论】:

  • 最好能展示您尝试过的代码。

标签: python pandas time-series


【解决方案1】:

您可以使用resample*2:

  • 首先重新采样到一个月 (M) 并获取差异 (max()-min())
  • 然后重新采样到 5 年 (5AS) 和 groupby 月并取 mean()

例如:

In []:
date_range = pd.date_range(start='2008-01-01',end='2017-12-31')
df = pd.DataFrame({'sale': np.random.randint(100, 200, size=date_range.size)},
                  index=date_range)

In []:
df1 = df.resample('M').apply(lambda g: g.max()-g.min())
df1.resample('5AS').apply(lambda g: g.groupby(g.index.month).mean()).unstack()

Out[]:
            sale                                                                  
              1     2     3     4     5     6     7     8     9     10    11    12
2008-01-01  95.4  90.2  95.2  95.4  93.2  93.8  91.8  95.6  93.4  93.4  94.2  93.8
2013-01-01  93.2  96.4  92.8  96.4  92.6  93.0  93.2  92.6  91.2  93.2  91.8  92.2

【讨论】:

    【解决方案2】:

    它可能看起来像这样(注意:没有累积销售价值)。这里的关键是执行一个 df.groupby() 传递 dt.year 和 dt.month。

    import pandas as pd
    import numpy as np
    
    df = pd.DataFrame({
        'date': pd.date_range(start='2016-01-01',end='2017-12-31'),
        'sale': np.random.randint(100,200, size = 365*2+1)
    })
    
    # Get month max, min and size (and as they are sorted - last and first)
    dfg = df.groupby([df.date.dt.year,df.date.dt.month])['sale'].agg(['last','first','size'])
    
    # Assign new cols (diff and avg) and drop max min size
    dfg = dfg.assign(diff = dfg['last'] - dfg['first'])
    dfg = dfg.assign(avg = dfg['diff'] / dfg['size']).drop(['last','first','size'], axis=1)
    
    # Rename index cols
    dfg.index = dfg.index.rename(['Year','Month'])
    
    print(dfg.head(6))
    

    返回:

                diff       avg
    Year Month                
    2016 1       -56 -1.806452
         2       -17 -0.586207
         3        30  0.967742
         4        34  1.133333
         5        46  1.483871
         6         2  0.066667
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2015-07-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-08-18
      • 1970-01-01
      • 2021-04-04
      • 2014-02-05
      相关资源
      最近更新 更多