【问题标题】:Downloaded PDF is corrupted, how can I download it correctly using Python?下载的 PDF 已损坏,如何使用 Python 正确下载?
【发布时间】:2020-06-23 06:26:59
【问题描述】:

我正在尝试使用 requests(python 2.7) 下载 pdf,这是我正在使用的代码:

file_resp = requests.post(file_url, data=payload, headers={"Referer":file_referer_url})
with open('test.pdf', 'wb') as f:
    f.write(file_resp.content)

但是,下载的 pdf 文件已损坏。这是我得到的响应标题:

Cache-Control: no-cache, no-store
Content-Encoding: gzip
Content-Type: text/html; charset=utf-8
Date: Tue, 23 Jun 2020 05:49:32 GMT
Expires: -1
Pragma: no-cache
Server: Microsoft-IIS/8.5
Transfer-Encoding: chunked
Vary: Accept-Encoding
X-AspNet-Version: 4.0.30319
x-frame-options: SAMEORIGIN
X-Powered-By: ASP.NET

另外,响应是这样的:
JVBERi0xLjQNCiW0tba3DQoxIDAgb2JqDQo8P...(像这样的长序列...)

请问,谁能帮我解决我可能做错的事情?

【问题讨论】:

  • 我可以看到 Content-Type 是“text/html;”,而不是 pdf...我不确定,这可能是下载时出现问题的原因。关于我应该如何处理这个问题的任何建议?
  • 虽然没有人帮助我完成这个小任务,但这里是链接,如果有人可能遇到类似情况,请通过参考解决它。 stackabuse.com/encoding-and-decoding-base64-strings-in-python

标签: python python-2.7 https python-requests


【解决方案1】:

响应是字符串类型,因此只需使用 base64.b64decode 对数据进行 Base64 解码并将其写入文件。

with open('application.pdf', 'wb') as file_to_save:
    decoded_pdf_data = base64.b64decode(file_resp.content)
    file_to_save.write(decoded_pdf_data)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-08-19
    • 1970-01-01
    • 2012-02-13
    • 1970-01-01
    相关资源
    最近更新 更多