【问题标题】:Download files from a list if not already downloaded如果尚未下载,则从列表中下载文件
【发布时间】:2010-07-04 01:05:15
【问题描述】:

我可以在 c# 中做到这一点,而且代码很长。

如果有人可以告诉我如何通过 python 完成这将是很酷的。

伪代码是:

url: www.example.com/somefolder/filename1.pdf

1. load file into an array (file contains a url on each line)
2. if file e.g. filename1.pdf doesn't exist, download file

脚本可以采用以下布局:

/python-downloader/
/python-downloader/dl.py
/python-downloader/urls.txt
/python-downloader/downloaded/filename1.pdf

【问题讨论】:

    标签: python


    【解决方案1】:

    这应该可以解决问题,尽管我假设 urls.txt 文件只包含 url。不是url: 前缀。

    import os
    import urllib
    
    DOWNLOADS_DIR = '/python-downloader/downloaded'
    
    # For every line in the file
    for url in open('urls.txt'):
        # Split on the rightmost / and take everything on the right side of that
        name = url.rsplit('/', 1)[-1]
    
        # Combine the name and the downloads directory to get the local filename
        filename = os.path.join(DOWNLOADS_DIR, name)
    
        # Download the file if it does not exist
        if not os.path.isfile(filename):
            urllib.urlretrieve(url, filename)
    

    【讨论】:

    • 哇,简直太简洁了!我乞求看看所有的炒作是关于什么的!谢谢大哥!
    • 使用 os.path.basename(url) 而不是在 '/' 上分割。
    【解决方案2】:

    这是 WoLpH 的 Python 3.3 脚本稍作修改的版本。

    #!/usr/bin/python3.3
    import os.path
    import urllib.request
    
    links = open('links.txt', 'r')
    for link in links:
        link = link.strip()
        name = link.rsplit('/', 1)[-1]
        filename = os.path.join('downloads', name)
    
        if not os.path.isfile(filename):
            print('Downloading: ' + filename)
            try:
                urllib.request.urlretrieve(link, filename)
            except Exception as inst:
                print(inst)
                print('  Encountered unknown error. Continuing.')
    

    【讨论】:

    • 您应该将 PATH 添加到脚本中。使用 PATH = "./downloads"
    【解决方案3】:

    Python 中的代码更少,你可以使用这样的代码:

    import urllib2
    improt os
    
    url="http://.../"
    # Translate url into a filename
    filename = url.split('/')[-1]
    
    if not os.path.exists(filename)
      outfile = open(filename, "w")
      outfile.write(urllib2.urlopen(url).read())
      outfile.close()
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2021-12-09
      • 1970-01-01
      • 2020-06-14
      • 2023-03-14
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多