【问题标题】:IndexError: index 1967 is out of bounds for axis 0 with size 1967IndexError:索引 1967 超出轴 0 的范围,大小为 1967
【发布时间】:2021-05-24 02:17:06
【问题描述】:

通过计算 p 值,我减少了大型稀疏文件中的特征数量。但我得到这个错误。我看过类似的帖子,但这段代码适用于非稀疏输入。你能帮忙吗? (如果需要我可以上传输入文件)

import statsmodels.formula.api as sm

def backwardElimination(x, Y, sl, columns):
    numVars = len(x[0])
    pvalue_removal_counter = 0

    for i in range(0, numVars):
        print(i, 'of', numVars)
        regressor_OLS = sm.OLS(Y, x).fit()
        maxVar = max(regressor_OLS.pvalues).astype(float)

        if maxVar > sl:
            for j in range(0, numVars - i):
                if (regressor_OLS.pvalues[j].astype(float) == maxVar):
                    x = np.delete(x, j, 1)
                    pvalue_removal_counter += 1
                    columns = np.delete(columns, j)

    regressor_OLS.summary()
    return x, columns

输出:

0 of 1970
1 of 1970
2 of 1970
Traceback (most recent call last):
  File "main.py", line 142, in <module>
    selected_columns)
  File "main.py", line 101, in backwardElimination
    if (regressor_OLS.pvalues[j].astype(float) == maxVar):
IndexError: index 1967 is out of bounds for axis 0 with size 1967

【问题讨论】:

    标签: python numpy statsmodels p-value index-error


    【解决方案1】:

    这是一个固定版本。

    我做了一些改变:

    1. 从 statsmodels.api 导入正确的OLS
    2. 在函数中生成columns
    3. 使用np.argmax查找最大值的位置
    4. 使用布尔索引来选择列。在伪代码中,它类似于 x[:, [True, False, True]],保留第 0 列和第 2 列。
    5. 如果没有东西可丢弃,请停止。
    import numpy as np
    # Wrong import. Not using the formula interface, so using statsmodels.api
    import statsmodels.api as sm
    
    def backwardElimination(x, Y, sl):
        numVars = x.shape[1]  # variables in columns
        columns = np.arange(numVars)
    
        for i in range(0, numVars):
            print(i, 'of', numVars)
            regressor_OLS = sm.OLS(Y, x).fit()
    
            if maxVar > sl:
                # Use boolean selection
                retain = np.ones(x.shape[1], bool)
                drop = np.argmax(regressor_OLS.pvalues)
                # Drop the highest pvalue(s)
                retain[drop] = False
                # Keep the x we with to retain
                x = x[:, retain]
                # Also keep their column indices
                columns = columns[retain]
            else:
                # Exit early if everything has pval above sl
                break
    
        # Show the final summary
        print(regressor_OLS.summary())
        return x, columns
    

    你可以测试一下

    x = np.random.standard_normal((1000,100))
    y = np.random.standard_normal(1000)
    backwardElimination(x,y,0.1)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2017-02-16
      • 2020-04-24
      • 2016-07-29
      • 2021-07-17
      • 2018-07-23
      • 2019-03-01
      • 2021-08-23
      • 2021-12-07
      相关资源
      最近更新 更多