【问题标题】:how to get a rolling mean with mean from previous window如何从上一个窗口中获得滚动平均值
【发布时间】:2022-01-17 21:48:45
【问题描述】:

我正在拼命寻找熊猫的解决方案。也许你能帮帮我。

我正在寻找一个考虑了先前平均值的滚动平均值。

df 看起来像这样:

index count
0 4
1 6
2 10
3 12

现在,使用rolling(window=2).mean() 函数我会得到这样的结果:

index count r_mean
0 4 NaN
1 6 5
2 10 8
3 12 11

我想考虑第一次计算的平均值,如下所示:

index count r_mean
0 4 NaN
1 6 5
2 10 7.5
3 12 9.5

在哪里,

row1: (4+6)/2=5

row2: (5+10)/2=7.5

row3: (7.5+12)/2=9.75

提前谢谢你!

【问题讨论】:

  • 很难矢量化

标签: python pandas mean rolling-computation


【解决方案1】:

编辑:实际上有在pandas中实现的方法ewm可以做这个计算

df['res'] = df['count'].ewm(alpha=0.5, adjust=False, min_periods=2).mean()

原答案:这是一种方法。因为一切都可以在系数为 2 的情况下发展。

# first create a series with power of 2
coef = pd.Series(2**np.arange(len(df)), df.index).clip(lower=2)

df['res'] = (coef.div(2)*df['count']).cumsum()/coef

print(df)
   index  count   res
0      0      4  2.00
1      1      6  5.00
2      2     10  7.50
3      3     12  9.75

如果需要,您可以使用df.loc[0, 'res'] = np.nan 屏蔽第一个值

【讨论】:

  • @ScottBoston 我意识到对于row2,您可以将计算重写为(2**0*4+2**0*6+2**1*10)/2**2,并将row3 重写为(2**0*4+2**0*6+2**1*10+ 2**2*12)/2**3,可以用幂2 进行推广:)
  • 令人着迷...不久前我一直在寻找这样的解决方案。 2系列的电源。当您沿系列滚动时减少影响。我喜欢它!
  • @ScottBoston 实际上使用 ewm 有一种更简单的方法;)
  • 哇......我需要把 ewm 放在口袋里。谢谢@BenT
  • 非常感谢@Ben.T! ewm直接帮了我..我花了多少时间,就这么简单!
【解决方案2】:

我们可以使用简单的python循环,如果你想加快速度,可以试试numba

l= []
n = 2
for x,y in zip(df['count'],df.index):
    try :
        l.append(np.nansum(x+l[y-n+1])/n)
    except:
        l.append(x)
df.loc[n-1:, 'new']=l[n-1:]
df
Out[332]: 
   index  count   new
0      0      4   NaN
1      1      6  5.00
2      2     10  7.50
3      3     12  9.75

【讨论】:

    猜你喜欢
    • 2018-05-11
    • 1970-01-01
    • 1970-01-01
    • 2020-07-12
    • 2020-08-29
    • 1970-01-01
    • 2018-10-03
    • 2020-04-23
    • 1970-01-01
    相关资源
    最近更新 更多