【问题标题】:See the score of each fold when cross validating a model using a for loop使用 for 循环交叉验证模型时查看每个折叠的分数
【发布时间】:2019-08-20 14:53:49
【问题描述】:

我想查看每个拟合模型的个人分数以可视化交叉验证的强度(我这样做是为了向我的同事展示为什么交叉验证很重要)。

我有一个 .csv 文件,其中包含 500 行、200 个自变量和 1 个二进制目标。我定义skf 使用StratifiedKFold 将数据折叠5 次。

我的代码如下所示:

X = data.iloc[0:500, 2:202]
y = data["target"]
skf = StratifiedKFold(n_splits = 5, random_state = 0)
clf = svm.SVC(kernel = "linear")
Scores = [0] * 5
for i, j in skf.split(X, y):
    X_train, y_train = X.iloc[i], y.iloc[i]
    X_test, y_test = X.iloc[j], y.iloc[j]
    clf.fit(X_train, y_train)
    clf.score(X_test, y_test)

如您所见,我为Scores 分配了一个包含 5 个零的列表。我想将 5 个预测中的每一个的 clf.score(X_test, y_test) 分配给列表。但是,索引 i 和 j 不是 {1, 2, 3, 4, 5}。相反,它们是用于折叠 X 和 y 数据帧的行号。

如何在此循环中将每个k 拟合模型的测试分数分配给Scores?我需要一个单独的索引吗?

我知道使用cross_val_score 确实可以做到所有这些,并为您提供k 分数的几何平均值。但是,我想向我的同事展示 sklearn 库中的交叉验证函数背后发生了什么。

提前致谢!

【问题讨论】:

    标签: python pandas for-loop cross-validation


    【解决方案1】:

    如果我理解了这个问题,并且您不需要任何特定的分数索引:

    from sklearn.model_selection import StratifiedKFold
    from sklearn.svm import SVC
    
    X = np.random.normal(size = (500, 200))
    y = np.random.randint(low = 0, high=2, size=500)
    skf = StratifiedKFold(n_splits = 5, random_state = 0)
    clf = SVC(kernel = "linear")
    Scores = []
    for i, j in skf.split(X, y):
        X_train, y_train = X[i], y[i]
        X_test, y_test = X[j], y[j]
        clf.fit(X_train, y_train)
        Scores.append(clf.score(X_test, y_test))
    

    结果是:

    >>>Scores
    [0.5247524752475248, 0.53, 0.5, 0.51, 0.4444444444444444]
    

    【讨论】:

    • 您正确理解了这个问题。这解决了我的问题。我总是忘记 Python 中 append 方法的存在。
    猜你喜欢
    • 2019-06-21
    • 2021-06-26
    • 2016-09-30
    • 2018-09-10
    • 1970-01-01
    • 2017-06-20
    • 2021-02-17
    • 1970-01-01
    • 2018-12-17
    相关资源
    最近更新 更多