【问题标题】:MLFlow RestException: RESOURCE_ALREADY_EXISTS error when starting runsMLFlow RestException:开始运行时出现 RESOURCE_ALREADY_EXISTS 错误
【发布时间】:2021-12-21 15:35:39
【问题描述】:

以前在 Azure 机器学习上使用 ML FLow 和 Databricks,从 9 月初开始使用 SKLearn 和 Stats 模型模型注册和跟踪模型超参数调整,没有任何问题。但是从大约 10 月 23 日开始,我开始收到这些错误:

RestException:RESOURCE_ALREADY_EXISTS:无法为实验 id=863468136127724、name=/my-experiment3、artifactLocation=dbfs:/databricks/mlflow-tracking/863468136127724 创建 AML 实验。现有的 AML 实验 id=c74bdea3-e382-4cdf-868a-ee1421de078e 和 name='/adb/5909321886823418/863468136127724/my-experiment3' 和 artifactLocation='' 不兼容。

Even when running a newly created experiment, it will throw this error

我们最近更新到 ml flow v1.21.0 但它似乎不是一个错误,因为 ML Flow github 上没有任何类似的东西,只是想知道是否有人遇到过类似的东西,因为我的想法用完了要查找的问题。

【问题讨论】:

  • 嗨,Rob,我也遇到了同样的错误……这太令人沮丧了,几周前可以运行的相同代码现在不行了。我向 Microsoft 团队提出了支持请求,如果我得到一些新的信息,我会通知你。你有关于这个问题的任何新信息吗?
  • 似乎对于我创建的每个实验,mlflow 还创建了一个 AML 实验,所有这些实验都指向同一个 artifactLocation=""。删除所有实验都没关系,垃圾收集器会检测到存在 artofactLocation="" 的实验,因此您尝试登录的任何新实验都会发生冲突。
  • 嗨,我们有同样的问题。您知道如何将 databricks 工作区与 azure ml 工作区分离(取消链接)吗?要链接它,可以从门户网站,但如何分离它?

标签: mlflow


【解决方案1】:

我联系了 Microsoft 支持团队。问题似乎是 azure databricks 错误地链接到 AML 工作区,它们提供了以下 ARM 模板,您应该在其中填写其中包含的参数:

{"$schema": "https://schema.management.azure.com/schemas/2015-01-01/deploymentTemplate.json#",
"contentVersion": "1.0.0.0",
"parameters": {
  "workspaceName": {
    "type": "string",
    "defaultValue": "",
    "metadata": {
        "description": "The name of the Azure Databricks workspace to create or update."
    }
  },
  "location": {
    "type": "string",
    "defaultValue": "[resourceGroup().location]",
    "metadata": {
      "description": "Location for all resources."
    }
  },
  "amlWorkspaceId": {
    "type": "string",
    "metadata": {
      "description": "The resource id of the Azure Machine Learning service workspace."
    },
    "defaultValue": ""
  }
},
"variables": {
  "managedResourceGroupName": "[concat('databricks-rg-', parameters('workspaceName'), '-', uniqueString(parameters('workspaceName'), resourceGroup().id))]"
},
"resources": [
  {
    "type": "Microsoft.Databricks/workspaces",
    "name": "[parameters('workspaceName')]",
    "location": "[parameters('location')]",
    "apiVersion": "2018-04-01",
    "properties": {
      "ManagedResourceGroupId": "[concat(subscription().id, '/resourceGroups/', variables('managedResourceGroupName'))]",
      "parameters": {
        "amlWorkspaceId": {
          "value": "[if(equals(parameters('amlWorkspaceId'),''), json('null'), parameters('amlWorkspaceId'))]"
        }
      }
    }
  }
],
"outputs": {
  "workspace": {
    "type": "object",
    "value": "[reference(resourceId('Microsoft.Databricks/workspaces', parameters('workspaceName')))]"
  }
}}

但是我无法在我的订阅中部署这个 ARM 模板,所以我决定按照这个线程 https://github.com/MicrosoftDocs/azure-docs/issues/80298 中的解决方案删除与我的 Databricks 自动链接的 Azure 机器学习服务,问题就解决了。

【讨论】:

    猜你喜欢
    • 2019-09-30
    • 1970-01-01
    • 1970-01-01
    • 2020-04-12
    • 2022-06-17
    • 1970-01-01
    • 2021-11-23
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多