【发布时间】:2020-03-22 22:48:35
【问题描述】:
根据AWS Glue文档,当连接类型为s3时,我们可以使用exlusions来排除文件:
https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect.html
“exclusions”:(可选)包含要排除的 Unix 样式 glob 模式的 JSON 列表的字符串。例如,"[\"**.pdf\"]" 排除所有 PDF 文件。有关 AWS Glue 支持的 glob 语法的更多信息,请参阅包含和排除模式。
我的 s3 存储桶喜欢关注,我想排除 test1 文件夹。
/mykkkkkk-test
test1/
testfolder/
11.json
22.json
test2/
1.json
test3/
2.json
test4/
3.json
test5/
4.json
我使用以下代码排除 test1 文件夹,但它仍然会在我的 test1 文件夹下使用 ETL 文件并且它不起作用
datasource0 = glueContext.create_dynamic_frame_from_options("s3",
{'paths': ["s3://mykkkkkk-test/"],
'exclusions': "[\"test1/**\"]",
'recurse':True,
'groupFiles': 'inPartition',
'groupSize': '1048576'},
format="json",
transformation_ctx = "datasource0")
exclusions 真的可以在 ETL pyspark 脚本中使用吗?我也尝试过,但没有任何效果
'exclusions': "[\"test1/**\"]",
'exclusions': ["test1/**"],
'exclusions': "[\"test1\"]",
【问题讨论】: