【问题标题】:aiobotocore-aiohttp - Get S3 file content and stream it in the responseaiobotocore-aiohttp - 获取 S3 文件内容并将其流式传输到响应中
【发布时间】:2017-03-05 08:34:30
【问题描述】:

我想使用 botocore 和 aiohttp 服务在 S3 上获取上传文件的内容。由于文件可能很大:

  • 我不想将整个文件内容存储在内存中,
  • 我希望在从 S3(aiobotocore、aiohttp)下载文件时能够处理其他请求,
  • 我希望能够对我下载的文件应用修改,所以我想逐行处理并将响应流式传输到客户端

目前,我的 aiohttp 处理程序中有以下代码:

import asyncio                                  
import aiobotocore                              

from aiohttp import web                         

@asyncio.coroutine                              
def handle_get_file(loop):                      

    session = aiobotocore.get_session(loop=loop)

    client = session.create_client(             
        service_name="s3",                      
        region_name="",                         
        aws_secret_access_key="",               
        aws_access_key_id="",                   
        endpoint_url="http://s3:5000"           
    )                                           

    response = yield from client.get_object(    
        Bucket="mybucket",                      
        Key="key",                              
    )                                           

每次我从给定文件中读取一行时,我都想发送响应。实际上, get_object() 返回一个带有 Body(ClientResponseContentProxy 对象)的字典。使用 read() 方法,我怎样才能获得一大块预期的响应并将其流式传输到客户端?

当我这样做时:

for content in response['Body'].read(10):
    print("----")                        
    print(content)          

循环内的代码永远不会执行。

但是当我这样做时:

result = yield from response['Body'].read(10)

我在结果中得到文件的内容。我对如何在这里使用 read() 有点困惑。

谢谢

【问题讨论】:

    标签: python amazon-s3 aiohttp botocore


    【解决方案1】:

    这是因为 aiobotocore api 与 botocore 的不同,这里 read() 返回一个 FlowControlStreamReader.read 生成器,您需要从中屈服

    看起来像这样(取自https://github.com/aio-libs/aiobotocore/pull/19)

    resp = yield from s3.get_object(Bucket='mybucket', Key='k')
    stream = resp['Body']
    try:
        chunk = yield from stream.read(10)
        while len(chunk) > 0:
          ...
          chunk = yield from stream.read(10)
    finally:
      stream.close()
    

    实际上在你的情况下你甚至可以使用readline()

    https://github.com/KeepSafe/aiohttp/blob/c39355bef6c08ded5c80e4b1887e9b922bdda6ef/aiohttp/streams.py#L587

    【讨论】:

    • 谢谢,这正是我需要的。
    猜你喜欢
    • 1970-01-01
    • 2011-01-25
    • 2018-02-19
    • 1970-01-01
    • 2020-11-28
    • 1970-01-01
    • 2014-03-19
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多