【问题标题】:Download big files via FTP with python使用python通过FTP下载大文件
【发布时间】:2012-01-09 13:13:24
【问题描述】:

我尝试每天从我的服务器下载一个备份文件到我的本地存储服务器,但我遇到了一些问题。

我写了这段代码(去掉了无用的部分,作为电子邮件功能):

import os
from time import strftime
from ftplib import FTP
import smtplib
from email.MIMEMultipart import MIMEMultipart
from email.MIMEBase import MIMEBase
from email.MIMEText import MIMEText
from email import Encoders

day = strftime("%d")
today = strftime("%d-%m-%Y")

link = FTP(ftphost)
link.login(passwd = ftp_pass, user = ftp_user)
link.cwd(file_path)
link.retrbinary('RETR ' + file_name, open('/var/backups/backup-%s.tgz' % today, 'wb').write)
link.delete(file_name) #delete the file from online server
link.close()
mail(user_mail, "Download database %s" % today, "Database sucessfully downloaded: %s" % file_name)
exit()

我使用如下 crontab 运行它:

40    23    *    *    *    python /usr/bin/backup-transfer.py >> /var/log/backup-transfer.log 2>&1

它适用于小文件,但它会冻结备份文件(大约 1.7Gb),下载的文件大约 1.2Gb 然后永远不会增长(我等了大约一天),并且日志文件是空的。

有什么想法吗?

ps:我使用的是 Python 2.6.5

【问题讨论】:

  • 要进一步解决问题,也许您可​​以使用来自FTP.retrbinarycallback 参数来收集有关下载进度的更多信息。此外,使用maxblocksize 可能会发现一些网络问题。

标签: python ftp file-transfer


【解决方案1】:

我用 ftplib 实现了代码,它可以监控连接、重新连接和重新下载文件以防失败。详情在这里:How to download big file in python via ftp (with monitoring & reconnect)?

【讨论】:

  • @roman 我正在使用你的脚本我正在下载一个大小为 2GB 的文件,因为我运行你的脚本时它在达到“dst_filesize”的大小时卡住了。
【解决方案2】:

对不起,如果我回答了我自己的问题,但我找到了解决方案。

我尝试了 ftputil 没有成功,所以我尝试了很多方法,最后,这行得通:

def ftp_connect(path):
    link = FTP(host = 'example.com', timeout = 5) #Keep low timeout
    link.login(passwd = 'ftppass', user = 'ftpuser')
    debug("%s - Connected to FTP" % strftime("%d-%m-%Y %H.%M"))
    link.cwd(path)
    return link

downloaded = open('/local/path/to/file.tgz', 'wb')

def debug(txt):
    print txt

link = ftp_connect(path)
file_size = link.size(filename)

max_attempts = 5 #I dont want death loops.

while file_size != downloaded.tell():
    try:
        debug("%s while > try, run retrbinary\n" % strftime("%d-%m-%Y %H.%M"))
        if downloaded.tell() != 0:
            link.retrbinary('RETR ' + filename, downloaded.write, downloaded.tell())
        else:
            link.retrbinary('RETR ' + filename, downloaded.write)
    except Exception as myerror:
        if max_attempts != 0:
            debug("%s while > except, something going wrong: %s\n \tfile lenght is: %i > %i\n" %
                (strftime("%d-%m-%Y %H.%M"), myerror, file_size, downloaded.tell())
            )
            link = ftp_connect(path)
            max_attempts -= 1
        else:
            break
debug("Done with file, attempt to download m5dsum")
[...]

在我的日志文件中我发现:

01-12-2011 23.30 - Connected to FTP
01-12-2011 23.30 while > try, run retrbinary
02-12-2011 00.31 while > except, something going wrong: timed out
    file lenght is: 1754695793 > 1754695793
02-12-2011 00.31 - Connected to FTP
Done with file, attempt to download m5dsum

遗憾的是,即使文件已完全下载,我也必须重新连接到 FTP,这在我的 cas 中没有问题,因为我也必须下载 md5sum。

如您所见,我无法检测到超时并重试连接,但是当我超时时,我只是重新连接;如果有人知道如何在不创建新的 ftplib.FTP 实例的情况下重新连接,请告诉我 ;)

【讨论】:

    【解决方案3】:

    您可以尝试设置超时。来自docs

    # timeout in seconds
    link = FTP(host=ftp_host, user=ftp_user, passwd=ftp_pass, acct='', timeout=3600)
    

    【讨论】:

      猜你喜欢
      • 2012-10-04
      • 2012-07-19
      • 1970-01-01
      • 1970-01-01
      • 2011-08-04
      • 2012-01-14
      • 2012-08-18
      • 2020-05-25
      相关资源
      最近更新 更多