【发布时间】:2018-01-10 20:26:59
【问题描述】:
我创建了一个 Pandas 数据框df:
df.head()
Out[1]:
A B DateTime
2010-01-01 50.662365 101.035099 2010-01-01
2010-01-02 47.652424 99.274288 2010-01-02
2010-01-03 51.387459 99.747135 2010-01-03
2010-01-04 52.344788 99.621896 2010-01-04
2010-01-05 47.106364 98.286224 2010-01-05
我可以添加 A 列的移动平均值:
df['A_moving_average'] = df.A.rolling(window=50, axis="rows") \
.apply(lambda x: np.mean(x))
问题:如何添加列 A 和 B的移动平均值?
这应该可以,但它给出了一个错误:
df['A_B_moving_average'] = df.rolling(window=50, axis="rows") \
.apply(lambda row: (np.mean(row.A) + np.mean(row.B)) / 2)
错误是:
NotImplementedError: ops for Rolling for this dtype datetime64[ns] are not implemented
附录 A:创建 Pandas 数据框的代码
这是我创建测试 Pandas 数据框df的方法:
import numpy.random as rnd
import pandas as pd
import numpy as np
count = 1000
dates = pd.date_range('1/1/2010', periods=count, freq='D')
df = pd.DataFrame(
{
'DateTime': dates,
'A': rnd.normal(50, 2, count), # Mean 50, standard deviation 2
'B': rnd.normal(100, 4, count) # Mean 100, standard deviation 4
}, index=dates
)
【问题讨论】:
-
更新:如果我删除
'DateTime': dates,行,则会出现不同的错误:AttributeError: 'numpy.ndarray' object has no attribute 'A'
标签: python pandas numpy time-series