【问题标题】:Why is statmotsmodels.Logit P > |z| different from scipy.stats pearsonr? They both are p_values, but they are giving me different values为什么是 statmotsmodels.Logit P > |z|不同于 scipy.stats pearsonr?它们都是 p_values,但它们给了我不同的值
【发布时间】:2021-02-20 22:02:38
【问题描述】:

这是我的 pandas 数据框和我的数据:

        c0  c1  c10 c11 c12 c13 c14 c15 c16 c3  c4  c5  c6  c7  c8  c9
index                                                               
0   1   49  2.0 0   2   2   0   1   6797.761892 130 269.0   0   1   163 0   0.0
1   0   61  0.0 1   2   2   1   3   4307.686943 138 166.0   0   0   125 1   3.6
2   0   46  0.0 2   3   2   0   1   4118.077502 140 311.0   0   1   120 1   1.8
3   0   69  1.0 3   3   2   1   0   7170.849469 140 254.0   0   0   146 0   2.0
4   0   51  1.0 0   2   2   1   0   5579.040145 100 222.0   0   1   143 1   1.2
... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ...
283 0   54  0.0 1   2   2   2   0   6293.123474 125 273.0   0   0   152 0   0.5
284 0   42  0.0 0   3   2   0   1   3303.841931 120 240.0   1   1   194 0   0.8
285 1   67  0.0 2   2   2   1   0   3383.029119 106 223.0   0   1   142 0   0.3
286 0   67  1.0 2   3   2   0   2   768.900795  125 254.0   1   1   163 0   0.2
287 0   60  0.0 1   3   2   0   0   1508.832825 130 253.0   0   1   144 1   1.4
288 rows × 16 columns

我使用 statsmodels 来获取 p_value:

log = sm.Logit(df['c0'], df.loc[:, df.columns != 'c0']).fit()
d1 = pd.DataFrame(index=log.pvalues.index, data=log.pvalues, columns=['statsmodels_pvalue'])

然后我也使用了 scipy 模块。 personr 函数返回相关性和 pvalue,如您所见,我正在附加返回 [1]。

index = []
output = []
for i in df.columns[1:]:
    index.append(i)
    output.append(pearsonr(df['c0'], df[i]) [1])    
d2 = pd.DataFrame(index=index, data=output, columns=['pearson_pvalue'])

pd.concat([d1,d2], axis=1)

结果:

    statsmodels_pvalue  pearson_pvalue
c1  0.155704            0.105977
c10 0.449688            0.697069
c11 0.041694            0.038457
c12 0.000269            0.000510
c13 0.012123            0.046765
c14 0.000114            0.000087
c15 0.587200            0.843444
c16 0.301656            0.025142
c3  0.434319            0.330075
c4  0.000163            0.000014
c5  0.792058            0.613432
c6  0.340877            0.454607
c7  0.843758            0.562002
c8  0.365109            0.030531
c9  0.238975            0.070500

【问题讨论】:

  • 你需要使用拦截
  • 关于personr函数?我该怎么做?

标签: python scipy statsmodels p-value pearson


【解决方案1】:

与统计相关的几点,也可以查看post like this。只有当您拟合截距并考虑 1 个因变量时,Pearson 和线性回归才等效。你正在做一个多元回归,这是行不通的。最后,你需要做一个普通的最小二乘,而不是逻辑回归。

下面将重现两者的 p 值:

from scipy.stats import pearsonr
import pandas as pd
import numpy as np
import statsmodels.api as sm

np.random.seed(111)

df = pd.DataFrame({'c0':np.random.uniform(0,1,50),
                   'c1':np.random.uniform(0,1,50),
                   'c2':np.random.uniform(0,1,50)})

variables = df.columns[1:]
output = []
for i in variables:
    lm = sm.OLS(df['c0'], sm.add_constant(df.loc[:, i])).fit()
    lm_p = lm.pvalues[1]
    pearson_p = pearsonr(df['c0'], df[i]) [1]
    output.append([lm_p,pearson_p])    

pd.DataFrame(output,index=variables,columns=['lm_p','pearson_p'])

    lm_p    pearson_p
c1  0.062513    0.062513
c2  0.781529    0.781529

【讨论】:

  • 就我而言,我的 Y 变量是二进制的(两个类),我如何将您的思路应用到我的问题上?我是否仍应将其视为数字(OLS)而不是两类问题? (逻辑)?
  • 嘿,你的问题是为什么 p 值不同,答案是因为你使用了不同的回归。如果您想找出哪个预测变量与您的 Y 变量相关联,那么您应该使用逻辑。希望这很清楚
猜你喜欢
  • 1970-01-01
  • 2013-08-18
  • 1970-01-01
  • 2021-09-07
  • 2015-07-14
  • 2020-05-22
  • 1970-01-01
  • 2010-11-25
  • 1970-01-01
相关资源
最近更新 更多