【问题标题】:exclusions doesn't work in AWS Glue ELT job s3 connection排除在 AWS Glue ELT 作业 s3 连接中不起作用
【发布时间】:2020-03-22 22:48:35
【问题描述】:

根据AWS Glue文档,当连接类型为s3时,我们可以使用exlusions来排除文件:

https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect.html

“exclusions”:(可选)包含要排除的 Unix 样式 glob 模式的 JSON 列表的字符串。例如,"[\"**.pdf\"]" 排除所有 PDF 文件。有关 AWS Glue 支持的 glob 语法的更多信息,请参阅包含和排除模式。

我的 s3 存储桶喜欢关注,我想排除 test1 文件夹。

/mykkkkkk-test
   test1/
      testfolder/
         11.json
         22.json
   test2/
      1.json
   test3/
      2.json
   test4/
      3.json
   test5/
      4.json

我使用以下代码排除 test1 文件夹,但它仍然会在我的 test1 文件夹下使用 ETL 文件并且它不起作用

datasource0 = glueContext.create_dynamic_frame_from_options("s3",
    {'paths': ["s3://mykkkkkk-test/"],
    'exclusions': "[\"test1/**\"]",
    'recurse':True,
    'groupFiles': 'inPartition',
    'groupSize': '1048576'}, 
    format="json",
    transformation_ctx = "datasource0")

exclusions 真的可以在 ETL pyspark 脚本中使用吗?我也尝试过,但没有任何效果

'exclusions': "[\"test1/**\"]",
'exclusions': ["test1/**"],
'exclusions': "[\"test1\"]",

【问题讨论】:

    标签: pyspark aws-glue


    【解决方案1】:

    尝试使用完整路径进行排除。

    datasource0 = glueContext.create_dynamic_frame.from_options(
    's3',
    {
        "paths": [
            's3://bucket/sample_data/'
        ],
        "recurse" : True,
        "exclusions" :  "[\"s3://bucket/sample_data/temp/**\"]"
    },
    "json",
    transformation_ctx = "datasource0")
    

    【讨论】:

      猜你喜欢
      • 2020-05-11
      • 1970-01-01
      • 2021-11-12
      • 2018-11-08
      • 1970-01-01
      • 1970-01-01
      • 2018-01-30
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多