【发布时间】:2019-11-08 11:38:49
【问题描述】:
我正在使用 Terraform 创建一个 EMR 集群 (emr-5.24.0),部署到一个私有子网中,其中包括 Spark、Hive 和 JupyterHub。
我在部署中添加了一个额外的配置 JSON,它应该将 Jupiter 笔记本的持久性添加到 S3(而不是本地磁盘上)。
整个架构包括一个到 S3 的 VPC 端点,我可以访问我尝试将笔记本写入到的存储桶。
配置集群时,JupyterHub 服务器无法启动。
登录到主节点并尝试为 jupyterhub 启动/重新启动 docker 容器没有帮助。
此持久性的配置如下所示:
[
{
"Classification": "jupyter-s3-conf",
"Properties": {
"s3.persistence.enabled": "true",
"s3.persistence.bucket": "${project}-${suffix}"
}
},
{
"Classification": "spark-env",
"Configurations": [
{
"Classification": "export",
"Properties": {
"PYSPARK_PYTHON": "/usr/bin/python3"
}
}
]
}
]
在 terraform EMR 资源定义中,会引用它:
configurations = "${data.template_file.configuration.rendered}"
这是读自:
data "template_file" "configuration" {
template = "${file("${path.module}/templates/cluster_configuration.json.tpl")}"
vars = {
project = "${var.project_name}"
suffix = "bucket"
}
}
当我不在笔记本上使用持久性时,一切正常,我可以登录 JupyterHub。
我很确定这不是 IAM 政策问题,因为 EMR 集群角色政策允许操作被定义为“s3:*”。
是否需要采取任何其他步骤才能使其发挥作用?
/K
【问题讨论】:
标签: amazon-s3 terraform amazon-emr terraform-provider-aws jupyterhub