【问题标题】:Pandas transformation of a column to row in a .rolling() fashionPandas 以 .rolling() 方式将列转换为行
【发布时间】:2021-09-26 03:14:33
【问题描述】:

假设我有以下系列:

inp = pd.Series(np.arange(10))

我想要做的是通过以下方式将其转换为np.array:

input 0 1 2 3 4
0 NaN NaN NaN Nan 0
1 NaN NaN NaN 0 1
2 NaN NaN 0 1 2
3 NaN 0 1 2 3
4 0 1 2 3 4
5 1 2 3 4 5
6 2 3 4 5 6

...等等。 输出中不应出现名为 input 的列,但我将其放在此处以使我的查询更加清晰。

我尝试的是以下内容:

matrix = [x.to_numpy() for x in list(inp.rolling(window=5, min_periods=5))]

问题是,我不能在matrix 上使用np.stack(),因为(即使我通过了min_periods=5)列表中每个项目的形状都不同。 我也觉得我忽略了一个非常简单的 pandas 命令:D。

非常感谢!

编辑: 我当前的解决方法是自定义函数。我想有比这个更好的解决方案:

def rolling_transform_series(x):
    length = len(x)
    array = []
    for idx in range(length): 
        s = x[idx-5:idx]
        if idx < 5:
            s = np.r_[np.zeros(5-idx), x[:idx]]
            s[s==0] = np.nan
        array.append(s)
    return np.array(array)

df = inp.apply(rolling_transform_series)

【问题讨论】:

    标签: python pandas dataframe numpy rolling-computation


    【解决方案1】:

    你可以用 shift 试试:

    import pandas as pd
    inp = pd.Series(np.arange(10))
    pd.DataFrame([inp.shift(s) for s in range(4,-6,-1)])
    
         0    1    2    3    4    5    6    7    8    9
    0  NaN  NaN  NaN  NaN  0.0  1.0  2.0  3.0  4.0  5.0
    1  NaN  NaN  NaN  0.0  1.0  2.0  3.0  4.0  5.0  6.0
    2  NaN  NaN  0.0  1.0  2.0  3.0  4.0  5.0  6.0  7.0
    3  NaN  0.0  1.0  2.0  3.0  4.0  5.0  6.0  7.0  8.0
    4  0.0  1.0  2.0  3.0  4.0  5.0  6.0  7.0  8.0  9.0
    5  1.0  2.0  3.0  4.0  5.0  6.0  7.0  8.0  9.0  NaN
    6  2.0  3.0  4.0  5.0  6.0  7.0  8.0  9.0  NaN  NaN
    7  3.0  4.0  5.0  6.0  7.0  8.0  9.0  NaN  NaN  NaN
    8  4.0  5.0  6.0  7.0  8.0  9.0  NaN  NaN  NaN  NaN
    9  5.0  6.0  7.0  8.0  9.0  NaN  NaN  NaN  NaN  NaN
    

    【讨论】:

    • 不幸的是,如果我理解正确的话,这在我的情况下并不适用。我的数据集大约有 800 万个数据点,我无法进行完整的转换并在之后对其进行剪切。但我会稍微修改一下 .shift() !谢谢安德烈亚斯
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2023-03-02
    • 1970-01-01
    • 2020-10-09
    相关资源
    最近更新 更多