【问题标题】:How to reduce the time to write pandas dataframes as table in Amazon Redshift如何减少在 Amazon Redshift 中将 pandas 数据帧写入表的时间
【发布时间】:2023-03-16 09:37:01
【问题描述】:

我正在使用这个在 Amazon Redshift 中编写 python pandas 数据框 -

df.to_sql('table_name', redshiftEngine, index = False, if_exists = 'replace' )

虽然我的数据框只有几千行和 50-100 列,但写一张表需要 15-20 分钟。我想知道这是否是 redshift 的正常表现?有什么办法可以优化这个过程,加快写表速度?

【问题讨论】:

    标签: python python-3.x pandas dataframe amazon-redshift


    【解决方案1】:

    更好的方法是使用 pandas 将数据帧存储为 CSV,将其上传到 S3 并使用 COPY 功能加载到 Redshift。这种方法甚至可以轻松处理数亿行。一般来说,Redshift 的写入性能不是很好——它是用来处理大量 ETL 操作(如 COPY)转储的数据负载。

    【讨论】:

    • single insert statement and commit,使用copy 导入数百万行,您将花费几乎相同的时间,加上@rvd 所说的,您甚至可以并行使用menifest 文件选项。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-01-29
    • 1970-01-01
    • 2021-03-02
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多