【问题标题】:Using the same preprocessing code for both training and inference in sagemaker在 sagemaker 中使用相同的预处理代码进行训练和推理
【发布时间】:2019-11-29 20:06:01
【问题描述】:

我正在为时间序列数据构建机器学习管道,其目标是频繁地重新训练和更新模型以进行预测。

  • 我编写了一个预处理代码,用于处理时间序列变量并对其进行转换。

我对如何在训练和推理中使用相同的预处理代码感到困惑?我应该编写一个 lambda 函数来预处理我的数据还是有其他方法

来源调查:

aws sagemaker 团队给出的两个示例使用 AWS Glue 进行 ETL 转换。

inference_pipeline_sparkml_xgboost_abalone

inference_pipeline_sparkml_blazingtext_dbpedia

我是 aws sagemaker 的新手,试图学习、理解和构建流程。任何帮助表示赞赏!

【问题讨论】:

  • 你的预处理代码在 scikit-learn 中吗?
  • 是的,还有 numpy、pandas 和 statsmodels。我尝试编写一个 lambda 来处理预处理,但对 lambda 层限制没有运气。
  • 这个答案能解决你的问题吗?

标签: machine-learning amazon-sagemaker amazon-machine-learning


【解决方案1】:

以倒退的方式回答问题。

从您的示例中,以下代码是两个模型放在一起的推理管道。在这里,我们需要删除 sparkml_model 并获取我们的 sklearn 模型。

sm_model = PipelineModel(name=model_name, role=role, models=[sparkml_model, xgb_model])

在放置 sklearn 模型之前,我们需要 SageMaker 版本的 SKLearn 模型。

首先使用 SageMaker Python 库创建 SKLearn Estimator。

sklearn_preprocessor = SKLearn(
    entry_point=script_path,
    role=role,
    train_instance_type="ml.c4.xlarge",
    sagemaker_session=sagemaker_session)

script_path - 这是包含所有预处理逻辑或转换逻辑的 python 代码。 'sklearn_abalone_featurizer.py' 在下面给出的链接中。

训练 SKLearn 估算器

sklearn_preprocessor.fit({'train': train_input})

从可以放入的 SKLearn Estimator 创建 SageMaker 模型 推理管道。

sklearn_inference_model = sklearn_preprocessor.create_model()

Inference PipeLineModel 创建将按如下所示进行修改。

sm_model = PipelineModel(name=model_name, role=role, models=[sklearn_inference_model, xgb_model])

更多详情,请参考以下链接。

https://github.com/awslabs/amazon-sagemaker-examples/blob/master/sagemaker-python-sdk/scikit_learn_inference_pipeline/Inference%20Pipeline%20with%20Scikit-learn%20and%20Linear%20Learner.ipynb

【讨论】:

    【解决方案2】:

    我在我的 python 脚本中使用了一个管道作为入口点。在此管道中,作为第一步,我正在执行预处理。管道保存为模型。因此,模型端点最终也包括预处理。详细信息(我使用的是 scikit,但对于 tensorflow 应该类似):

    如果你想像这样打电话给你的火车:

    from sagemaker.sklearn.estimator import SKLearn
    
    sklearn_estimator = SKLearn(
      entry_point='script.py',
      role = 'xxx',
      train_instance_count=1,
      train_instance_type='ml.c5.xlarge',
      framework_version='0.20.0',
      hyperparameters = {'cross-validation': 5,
                       'scoring': 'accuracy'})
    

    那么你就有了一个入口点脚本。在这个脚本('script.py')中,您可以有几个步骤成为最终保存的模型的一部分。例如:

    tfidf = TfidfVectorizer(strip_accents=None,
                        lowercase=False,
                        preprocessor=None)
    
    ....
    
    lr_tfidf = Pipeline([('vect', tfidf),
                     ('clf', LogisticRegression(random_state=0))])
    

    您需要在训练结束后通过 joblib.dump 保存您的模型。此存储模型用于创建 sagemaker 模型和模型端点。当我最终调用 predictor.predict(X_test) 时,管道的第一步(我的 proeprocessing)也被执行并应用于 X_test。

    Sagemakers 支持不同的预处理方式。我只是想分享一个相当简单的,适合我的场景。我正在使用 GridSearch 作为 script.py 中管道步骤的参数。

    【讨论】:

    • 这里的Pipeline函数是什么。您是说 PipelineModel 还是其他功能?
    • 刚刚实现了它的 sklearn 功能,以防其他人想知道。见这里:scikit-learn.org/stable/modules/generated/…
    • @AndiSchroff 这是一个 SKLearn 管道,而不是 SageMaker 推理管道...
    猜你喜欢
    • 2019-06-20
    • 1970-01-01
    • 2019-08-09
    • 2014-12-19
    • 2018-11-29
    • 1970-01-01
    • 1970-01-01
    • 2018-05-07
    • 2021-02-04
    相关资源
    最近更新 更多