【发布时间】:2021-05-01 12:15:36
【问题描述】:
我有以下数据,我想将使用样条的插值方法应用于最后 4 个数字(我知道这是外推):
import numpy as np
x = [
18.792571,
19.170139,
19.370556,
19.393820,
19.239932,
18.908891,
18.400699,
17.892507,
17.384314,
16.876122,
16.367930,
15.859737,
np.nan,
np.nan,
np.nan,
np.nan
]
我正在运行 pandas interpolate 并且发生了一件非常奇怪的事情,就像代码一样
import pandas as pd
pd.Series(x).interpolate(
method="spline",
order=1
)
返回
0 18.792571
1 19.170139
2 19.370556
3 19.393820
4 19.239932
5 18.908891
6 18.400699
7 17.892507
8 17.384314
9 16.876122
10 16.367930
11 15.859737
12 16.103099
13 15.790022
14 15.476945
15 15.163868
dtype: float64
因此,虽然数据的趋势显然是负面的,但因为很早的指数,插值产生了向上的跳跃。使用 scipy 运行相同的计算时
import scipy.interpolate as inp
train_x = [_ for _ in x if _ > 0]
s = inp.InterpolatedUnivariateSpline(range(len(train_x)), train_x, k=1)
ynew = s(range(len(x)))
ynew[12:]
我明白了
array([15.351544, 14.843351, 14.335158, 13.826965])
在这种情况下,插值没有向上变化,所以结果对我来说是有意义的。
那么我的问题是:
- 为什么 pandas 和 scipy 的结果不一样?
- 如何让 pandas
interpolate给出我使用 scipy 获得的结果? - 为什么熊猫会发生这种向上的变化?
提前致谢!
编辑
使用 scipy interp1d 我有同样的问题:
s = inp.interp1d(range(len(train_x)), train_x, kind=1, fill_value='extrapolate')
ynew = s(range(len(x)))
ynew[12:]
给予
array([15.351544, 14.843351, 14.335158, 13.826965])
【问题讨论】:
-
嗯,这很奇怪。有趣的是,Pandas
interpolate和method=spline的输出实际上与使用线性回归在整个数据集上推断的数据相匹配。如果您使用不同的方法(例如method=slinear),您将获得与直接 Scipy 实现类似的结果。我实际上不清楚 Pandas 如何解释method=spline,即它执行的确切 scipy 函数和参数。
标签: python pandas numpy scipy interpolation