前面的答案是一个很好的基本开始,但我想实现下面所述的高级目标。总的来说,我觉得awswrangler 是要走的路。
- 读取.gzip
- 只读取前 5 行而不下载完整文件
- 明确传递凭据(确保您没有将它们提交给代码!!)
- 使用完整的 s3 路径
以下是一些有效的方法
import boto3
import pandas as pd
import awswrangler as wr
boto3_creds = dict(region_name="us-east-1", aws_access_key_id='', aws_secret_access_key='')
boto3.setup_default_session(**boto3_creds)
s3 = boto3.client('s3')
# read first 5 lines from file path
obj = s3.get_object(Bucket='bucket', Key='path.csv.gz')
df = pd.read_csv(obj['Body'], nrows=5, compression='gzip')
# read first 5 lines from directory
dft_xp = pd.concat(list(wr.s3.read_csv(wr.s3.list_objects('s3://bucket/path/')[0], chunksize=5, nrows=5, compression='gzip')))
# read all files into pandas
df_xp = wr.s3.read_csv(wr.s3.list_objects('s3://bucket/path/'), compression='gzip')
没有使用 s3fs 不确定是否使用 boto3。
对于使用 dask 的分布式计算,这可行,但它使用 s3fs afaik 并且显然 gzip 无法并行化。
import dask.dataframe as dd
dd.read_csv('s3://bucket/path/*', storage_options={'key':'', 'secret':''}, compression='gzip').head(5)
dd.read_csv('s3://bucket/path/*', storage_options={'key':'', 'secret':''}, compression='gzip')
# Warning gzip compression does not support breaking apart files Please ensure that each individual file can fit in memory