【问题标题】:cublas batched gemm throw not supported error with large batch sizecublas batched gemm throw not supported error with large batch size
【发布时间】:2019-01-01 04:02:52
【问题描述】:

我正在调用 cublasGemmStridedBatchedEx() API。我让第一个矩阵大步前进,第二个矩阵固定。该程序适用于小输入,但在大批量时会引发 CUBLAS_STATUS_NOT_SUPPORTED 错误。

根据cublas documentation,这意味着不支持数据类型或算法。我看不出增加批量大小如何改变数据类型。我使用默认的启发式 GEMM 算法。

我使用 CUDA9.2 编译代码并在 GTX 1050 卡上运行。

代码:

#include <stdio.h>
#include <stdlib.h>

#include <cuda_runtime.h>
#include <cublas_v2.h>
#include <cuda_fp16.h>

#include "nvidia_helper/checkCudaErrors.h"

int UPPER_BOUND = 4096;

int main() {
    half* F4_re;
    half* X_split;
    float* result1;

    int M = 16;
    int B = 256*64;

    checkCudaErrors(cudaMallocManaged((void **) &F4_re, 4 * 4 * sizeof(half)));
    checkCudaErrors(cudaMallocManaged((void **) &X_split, M * 4 * B * 4 * sizeof(half)));
    checkCudaErrors(cudaMallocManaged((void **) &result1, M * 4 * B * 4 * sizeof(float)));

    F4_re[0] = 1.0f;
    F4_re[1] = 1.0f;
    F4_re[2] = 1.0f;
    F4_re[3] = 1.0f;
    F4_re[4] = 1.0f;
    F4_re[5] = 0.0f;
    F4_re[6] =-1.0f;
    F4_re[7] = 0.0f;
    F4_re[8] = 1.0f;
    F4_re[9] =-1.0f;
    F4_re[10] = 1.0f;
    F4_re[11] =-1.0f;
    F4_re[12] = 1.0f;
    F4_re[13] = 0.0f;
    F4_re[14] =-1.0f;
    F4_re[15] = 0.0f;

    srand(time(NULL));
    for (int i = 0; i < M * 4 * B * 4; i++) {
       X_split[i] = (float)rand() / (float)(RAND_MAX) * 2 * UPPER_BOUND - UPPER_BOUND;
    }

    cublasStatus_t status;
    cublasHandle_t handle;
    float alpha = 1.0f, beta = 0.0f; 

    status = cublasCreate(&handle);
    if (status != CUBLAS_STATUS_SUCCESS) {
        fprintf(stderr, "!!!! CUBLAS initialization error\n");
        exit(1);
    }
    status = cublasSetMathMode(handle, CUBLAS_TENSOR_OP_MATH); // allow Tensor Core
    if (status != CUBLAS_STATUS_SUCCESS) {
        fprintf(stderr, "!!!! CUBLAS setting math mode error\n");
        exit(1);
    }

    long long int stride = M * 4;

    status = cublasGemmStridedBatchedEx(handle, CUBLAS_OP_N, CUBLAS_OP_N, M, 4, 4, &alpha, X_split,
             CUDA_R_16F, M, stride, F4_re, CUDA_R_16F, 4, 0, &beta, result1, CUDA_R_32F, M, stride, B * 4, CUDA_R_32F, CUBLAS_GEMM_DEFAULT);
    if (status != CUBLAS_STATUS_SUCCESS) {
        fprintf(stderr, "!!!! CUBLAS kernel execution error: %d .\n", status);
        exit(1);
    }

    status = cublasDestroy(handle);
    if (status != CUBLAS_STATUS_SUCCESS) {
        fprintf(stderr, "!!!! shutdown error (A)\n");
        exit(1);
    }

    checkCudaErrors(cudaFree(F4_re));
    checkCudaErrors(cudaFree(X_split));
    checkCudaErrors(cudaFree(result1));

    return 0;
}

【问题讨论】:

  • 你应该提供一个minimal reproducible example,见第1项here你提供的不是一个完整的例子。
  • @RobertCrovella 描述已更新
  • @XiaoheCheng 请在代码停止工作的特定矩阵大小处添加信息。这个批处理 API 是为许多小矩阵的用例设计的(通过经典的 BLAS API 调用处理效率低下)并且它可能(推测!)对矩阵大小(可能是最大维度)有限制大约 150-200) 被无意中从文档中遗漏了。
  • 根据我的测试,如果传递给 cublasGemmStridedBatchedEx 的批次计数为 65535 或更大,则提供的代码将失败(返回 cublas 错误代码 15 - CUBLAS_STATUS_NOT_SUPPORTED)并具有给定的矩阵尺寸。需要明确的是,我的观察:65536 失败,65535 失败,65534 通过。
  • 我已经向 NVIDIA 提交了一个内部错误。我目前没有任何进一步的信息。

标签: cuda precision cublas


【解决方案1】:

这个问题应该在刚刚发布的 CUBLAS 补丁版本中得到纠正,可用here

寻找这个补丁:

补丁 1(2018 年 8 月 6 日发布)

这是一个必须安装在正确安装 CUDA 9.2.148 之上的补丁。

首先您必须安装 CUDA 9.2.148。然后安装补丁。

【讨论】:

    猜你喜欢
    • 2023-03-18
    • 2017-12-20
    • 2018-07-28
    • 1970-01-01
    • 1970-01-01
    • 2017-09-22
    • 1970-01-01
    • 2018-07-09
    • 2018-01-07
    相关资源
    最近更新 更多