【问题标题】:Python seek on remote file using HTTPPython 使用 HTTP 查找远程文件
【发布时间】:2009-12-28 19:57:07
【问题描述】:

如何查找远程 (HTTP) 文件上的特定位置,以便仅下载该部分?

假设远程文件上的字节为:1234567890

我想寻找 4 并从那里下载 3 个字节,所以我会得到:456

还有,我如何检查远程文件是否存在? 我试过了,os.path.isfile() 但是当我传递远程文件 url 时它返回 False。

【问题讨论】:

  • “远程”是什么意思?
  • 你使用什么协议? HTTP? FTP? NFS? SFTP?

标签: python http seek


【解决方案1】:

如果是通过HTTP方式下载远程文件,需要设置Range标头。

检查in this example 是如何做到的。看起来像这样:

myUrlclass.addheader("Range","bytes=%s-" % (existSize))

编辑:I just found a better implementation。这个类使用起来非常简单,可以在文档字符串中看到。

class HTTPRangeHandler(urllib2.BaseHandler):
"""Handler that enables HTTP Range headers.

This was extremely simple. The Range header is a HTTP feature to
begin with so all this class does is tell urllib2 that the 
"206 Partial Content" reponse from the HTTP server is what we 
expected.

Example:
    import urllib2
    import byterange

    range_handler = range.HTTPRangeHandler()
    opener = urllib2.build_opener(range_handler)

    # install it
    urllib2.install_opener(opener)

    # create Request and set Range header
    req = urllib2.Request('http://www.python.org/')
    req.header['Range'] = 'bytes=30-50'
    f = urllib2.urlopen(req)
"""

def http_error_206(self, req, fp, code, msg, hdrs):
    # 206 Partial Content Response
    r = urllib.addinfourl(fp, hdrs, req.get_full_url())
    r.code = code
    r.msg = msg
    return r

def http_error_416(self, req, fp, code, msg, hdrs):
    # HTTP's Range Not Satisfiable error
    raise RangeError('Requested Range Not Satisfiable')

更新:“更好的实现”已移至byterange.py 文件中的github: excid3/urlgrabber。

【讨论】:

    【解决方案2】:

    我强烈推荐使用requests 库。它很容易成为我用过的最好的 HTTP 库。特别是,要完成您所描述的,您将执行以下操作:

    import requests
    
    url = "http://www.sffaudio.com/podcasts/ShellGameByPhilipK.Dick.pdf"
    
    # Retrieve bytes between offsets 3 and 5 (inclusive).
    r = requests.get(url, headers={"range": "bytes=3-5"})
    
    # If a 4XX client error or a 5XX server error is encountered, we raise it.
    r.raise_for_status()
    

    【讨论】:

    • 当时没有 requests 库,但是现在这让事情变得更简单了。
    【解决方案3】:

    AFAIK,使用 fseek() 或类似方法是不可能的。您需要使用 HTTP Range 标头来实现此目的。服务器可能支持也可能不支持此标头,因此您的里程可能会有所不同。

    import urllib2
    
    myHeaders = {'Range':'bytes=0-9'}
    
    req = urllib2.Request('http://www.promotionalpromos.com/mirrors/gnu/gnu/bash/bash-1.14.3-1.14.4.diff.gz',headers=myHeaders)
    
    partialFile = urllib2.urlopen(req)
    
    s2 = (partialFile.read())
    

    编辑:这当然是假设远程文件是指存储在 HTTP 服务器上的文件...

    如果你想要的文件在 FTP 服务器上,FTP 只允许指定一个 start 偏移量而不是一个范围。如果这是你想要的,那么下面的代码应该可以做到(未经测试!)

    import ftplib
    fileToRetrieve = 'somefile.zip'
    fromByte = 15
    ftp = ftplib.FTP('ftp.someplace.net')
    outFile = open('partialFile', 'wb')
    ftp.retrbinary('RETR '+ fileToRetrieve, outFile.write, rest=str(fromByte))
    outFile.close()
    

    【讨论】:

    • 您还应该处理 206 响应代码,因为如果您使用的是 HTTP 范围标头,它们可能是可以接受的。
    • 很公平。不过你的回答是这样的:)
    【解决方案4】:

    您可以使用httpio 访问远程 HTTP 文件,就像它们是本地文件一样:

    pip install httpio
    
    import zipfile
    import httpio
    
    url = "http://some/large/file.zip"
    with httpio.open(url) as fp:
        zf = zipfile.ZipFile(fp)
        print(zf.namelist())
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-11-25
      • 2021-03-25
      • 2012-04-06
      • 1970-01-01
      • 1970-01-01
      • 2021-08-06
      • 1970-01-01
      • 2012-03-24
      相关资源
      最近更新 更多