我今天遇到了一个有趣的角落案例。
如果我们只查看非常少量的样本,Spearman 和 Pearson 之间的差异可能会非常大。
在以下情况下,这两种方法报告了完全相反的相关性。
一些快速的经验法则来决定 Spearman 与 Pearson:
- Pearson 的假设是恒定的方差和线性(或相当接近的东西),如果不满足这些假设,可能值得尝试 Spearman。
- 上面的例子是一个极端情况,只有在有少数 (100 个数据点,并且数据呈线性或接近线性,则 Pearson 将与 Spearman 非常相似。
- 如果您认为线性回归是分析数据的合适方法,那么 Pearson 的输出将匹配线性回归斜率的符号和大小(如果变量已标准化)。
- 如果您的数据包含一些线性回归无法识别的非线性成分,则首先尝试通过应用转换(可能是 log e)将数据整理成线性形式。如果这不起作用,那么 Spearman 可能是合适的。
- 我总是先尝试 Pearson 的,如果不行,我会尝试 Spearman。
- 您能否添加更多经验法则或更正我刚刚推断出的那些?我已将此问题设为社区 Wiki,因此您可以这样做。
附言这是重现上图的 R 代码:
# Script that shows that in some corner cases, the reported correlation for spearman can be
# exactly opposite to that for pearson. In this case, spearman is +0.4 and pearson is -0.4.
y = c(+2.5,-0.5, -0.8, -1)
x = c(+0.2,-3, -2.5,+0.6)
plot(y ~ x,xlim=c(-6,+6),ylim=c(-1,+2.5))
title("Correlation: corner case for Spearman vs. Pearson\nNote that they are exactly opposite each other (-0.4 vs. +0.4)")
abline(v=0)
abline(h=0)
lm1=lm(y ~ x)
abline(lm1,col="red")
spearman = cor(y,x,method="spearman")
pearson = cor(y,x,method="pearson")
legend("topleft",
c("Red line: regression.",
sprintf("Spearman: %.5f",spearman),
sprintf("Pearson: +%.5f",pearson)
))