【问题标题】:downloading multiple files with urllib.urlretrieve使用 urllib.urlretrieve 下载多个文件
【发布时间】:2015-10-16 02:12:42
【问题描述】:

我正在尝试从网站下载多个文件。 网址类似于:foo.com/foo-1.pdf。 由于我希望将这些文件存储在我选择的目录中, 我写了以下代码:

import os
from urllib import urlretrieve
ext = ".pdf"
for i in range(1,37):
    print "fetching file " + str(i)
    url = "http://foo.com/Lec-" + str(i) + ext
    myPath = "/dir/"
    filename = "Lec-"+str(i)+ext
    fullfilename = os.path.join(myPath, filename)
    x = urlretrieve(url, fullfilename)

编辑:完整的错误信息。

Traceback (most recent call last):
File "scraper.py", line 10, in <module>
x = urlretrieve(url, fullfilename)
File "/usr/lib/python2.7/urllib.py", line 94, in urlretrieve
return _urlopener.retrieve(url, filename, reporthook, data)
File "/usr/lib/python2.7/urllib.py", line 244, in retrieve
tfp = open(filename, 'wb')
IOError: [Errno 2] No such file or directory: /dir/Lec-1.pdf'

如果有人能指出我哪里出错了,我将不胜感激。

提前致谢!

【问题讨论】:

  • 你说应该下载到/dir/。这个目录存在吗?
  • 我又查了一下,确实存在。
  • 1.您可以发布整个错误消息吗? 2.你不需要urlopen(url)。
  • 表示目录不存在。尝试使用import os 和os.makedirs('/dir/', exist_ok=True) 在最后一行之前创建它。

标签: python-2.7 urllib python-2.x


【解决方案1】:

对我而言,您的代码有效(Python3.9)。因此,请确保您的脚本可以访问您指定的目录。此外,您似乎正在尝试打开一个不存在的文件。因此,请确保在打开文件之前已下载该文件:

fullfilename = os.path.abspath("d:/DownloadedFiles/Lec-1.pdf")
print(fullfilename)
if os.path.exists(fullfilename): # open file only if it exists
    with open(fullfilename, 'rb') as file:
        content = file.read() # read file's content
        print(content[:150])  # print only the first 150 characters

输出如下:

C:/Users/Administrator/PycharmProjects/Tests/dtest.py
d:\DownloadedFiles\Lec-1.pdf
b'%PDF-1.6\r%\xe2\xe3\xcf\xd3\r\n2346 0 obj <</Linearized 1/L 1916277/O 2349/E 70472/N 160/T 1869308/H [ 536 3620]>>\rendobj\r       \r\nxref\r\n2346 12\r\n0000000016 00000 n\r'

Process finished with exit code 0

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2017-10-19
    • 2018-05-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多