【发布时间】:2020-03-18 16:27:53
【问题描述】:
我能够与我的 Databricks FileStore DBFS 建立连接并访问文件存储。
可以使用 Pyspark 读取、写入和转换数据,但是当我尝试使用本地 Python API(例如 pathlib 或 OS 模块)时,我无法通过第一级 DBFS 文件系统
我可以使用魔法命令:
%fs ls dbfs:\mnt\my_fs\... 完美运行并列出所有子目录?
但如果我这样做 os.listdir('\dbfs\mnt\my_fs\') 它会返回 ['mount.err'] 作为返回值
我在新集群上测试过,结果是一样的
我在 Databricks Runtine 6.1 版和 Apache Spark 2.4.4 上使用 Python
有没有人可以提供建议。
编辑:
连接脚本:
我使用 Databricks CLI 库来存储我的凭据,这些凭据根据 databricks 文档进行格式化:
def initialise_connection(secrets_func):
configs = secrets_func()
# Check if the mount exists
bMountExists = False
for item in dbutils.fs.ls("/mnt/"):
if str(item.name) == r"WFM/":
bMountExists = True
# drop if exists to refresh credentials
if bMountExists:
dbutils.fs.unmount("/mnt/WFM")
bMountExists = False
# Mount a drive
if not (bMountExists):
dbutils.fs.mount(
source="adl://test.azuredatalakestore.net/WFM",
mount_point="/mnt/WFM",
extra_configs=configs
)
print("Drive mounted")
else:
print("Drive already mounted")
【问题讨论】:
标签: python azure databricks azure-databricks