在 cmets 中组织我们冗长的讨论以获得答案。
矩阵乘法运算符/函数,如"%*%",crossprod,tcrossprod` 需要具有“数字”、“复杂”或“逻辑”模式的矩阵。但是,您的矩阵具有“字符”模式。
library(mlbench)
data(BreastCancer)
X <- as.matrix(BreastCancer[, 1:10])
mode(X)
#[1] "character"
您可能会感到惊讶,因为数据集似乎包含数字数据:
head(BreastCancer[, 1:10])
# Id Cl.thickness Cell.size Cell.shape Marg.adhesion Epith.c.size
#1 1000025 5 1 1 1 2
#2 1002945 5 4 4 5 7
#3 1015425 3 1 1 1 2
#4 1016277 6 8 8 1 3
#5 1017023 4 1 1 3 2
#6 1017122 8 10 10 8 7
# Bare.nuclei Bl.cromatin Normal.nucleoli Mitoses
#1 1 3 1 1
#2 10 3 2 1
#3 2 3 1 1
#4 4 3 7 1
#5 1 3 1 1
#6 10 9 7 1
但是你被印刷风格误导了。 这些列实际上是字符或因素:
lapply(BreastCancer[, 1:10], class)
#$Id
#[1] "character"
#
#$Cl.thickness
#[1] "ordered" "factor"
#
#$Cell.size
#[1] "ordered" "factor"
#
#$Cell.shape
#[1] "ordered" "factor"
#
#$Marg.adhesion
#[1] "ordered" "factor"
#
#$Epith.c.size
#[1] "ordered" "factor"
#
#$Bare.nuclei
#[1] "factor"
#
#$Bl.cromatin
#[1] "factor"
#
#$Normal.nucleoli
#[1] "factor"
#
#$Mitoses
#[1] "factor"
当您执行as.matrix 时,这些列都被强制转换为“字符”(有关详细说明,请参阅R: Why am I not getting type or class "factor" after converting columns to factor?)。
所以要进行矩阵乘法,我们需要正确地将这些列强制转换为“数字”。
dat <- BreastCancer[, 1:10]
## character to numeric
dat[[1]] <- as.numeric(dat[[1]])
## factor to numeric
dat[2:10] <- lapply( dat[2:10], function (x) as.numeric(levels(x))[x] )
## get the matrix
X <- data.matrix(dat)
mode(X)
#[1] "numeric"
现在您可以进行矩阵向量乘法等操作了。
## some possible matrix-vector multiplications
beta <- runif(10)
yhat <- X %*% beta
## add prediction back to data frame
dat$prediction <- yhat
但是,我怀疑这是为您的逻辑回归模型获取预测值的正确方法,因为当您使用因子构建模型时,模型矩阵不是上面的X,而是一个虚拟矩阵。我强烈推荐你使用predict。
这条线也适用于我:as.matrix(sapply(dat, as.numeric))
看来你很幸运。数据集恰好具有与数值相同的因子水平。一般来说,将因子转换为数字应该使用我所做的方法。比较
f <- gl(4, 2, labels = c(12.3, 0.5, 2.9, -11.1))
#[1] 12.3 12.3 0.5 0.5 2.9 2.9 -11.1 -11.1
#Levels: 12.3 0.5 2.9 -11.1
as.numeric(f)
#[1] 1 1 2 2 3 3 4 4
as.numeric(levels(f))[f]
#[1] 12.3 12.3 0.5 0.5 2.9 2.9 -11.1 -11.1
文档页面?factor对此进行了介绍。