【问题标题】:Testing the normality and correlation of the feature and label values测试特征和标签值的正态性和相关性
【发布时间】:2022-01-24 12:17:30
【问题描述】:

我有一个数据集,它存储在一个 2D numpy 数组中。我想测试作为数组列的每个特征的正态性和相关性,然后绘制它。

我知道使用R,可以通过运行以下命令轻松完成:

shapiro.test(Class$Feature)
ggqqplot(Wage$age, ylab = "Feature")

同样,在 R 中,可以通过运行以下命令轻松完成相关性测试:

res <- cor.test(Class$Feature, Class$class, method = "pearson")

如何在 python 中执行这些步骤?

我用下面的多列数据集尝试了Scipy的Normaltest,但id没有用。

from scipy import stats
df = pd.DataFrame(data)
k2, p = stats.normaltest(df[:,1], df[:,5]) #Testing Feature 1 agains Feature 5
print (p)

【问题讨论】:

  • 关于 Pearson,使用 corrcoef(例如来自 numpy)有效吗?

标签: python r numpy scipy data-visualization


【解决方案1】:

经过大量搜索后,我注意到使用numpy 数组可能不是解决此问题的合适方法。这就是为什么我将我的数据集加载到 pandas 数据框中,然后使用以下代码:

from scipy.stats import shapiro
import pylab
import scipy.stats as stats
def test_normality(data_frame, features, feature_for_test):
    for feature in features:
        print("Test Result: " + str(shapiro(data_frame[feature])))
        stats.probplot(data_frame[feature], dist="norm", plot=pylab)
        pylab.show()

test_normality(data_frame, ["feature1","feature2", "feature3"], "feature_for_test")

对于相关性测试,我使用了以下代码:

from scipy.stats import pearsonr
def correlation_test(data_frame, features, feature_for_test):
for feature in features:
    cor, _ = pearsonr(data_frame[feature], data_frame[feature_for_test])
    print("Pearson Correlation Test Result: %.3f" % cor)

correlation_test(data_frame, ["feature1","feature2", "feature3"], "feature_for_test")

【讨论】:

    猜你喜欢
    • 2022-07-05
    • 2019-04-25
    • 2019-03-09
    • 1970-01-01
    • 1970-01-01
    • 2017-01-02
    • 2011-07-16
    • 2018-05-20
    • 1970-01-01
    相关资源
    最近更新 更多