【问题标题】:How to solve the issue of the downloaded Excel files presenting the error?如何解决下载的Excel文件出现错误的问题?
【发布时间】:2022-06-14 23:01:09
【问题描述】:

我正在尝试从代码中看到的链接下载文件。但是,当我打开下载的文件时出现以下错误。我该如何解决这个问题?

请看下面的代码:

import os
import requests
from bs4 import BeautifulSoup
# Python 3.x
from urllib.request import urlopen, urlretrieve, quote
from urllib.parse import urljoin
import urllib

headers={"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.71 Safari/537.36"}
resp = requests.get("https://www.elections.on.ca/en/resource-centre/elections-results.html#accordion2022ge")
soup = BeautifulSoup(resp.text,"html.parser")

for link in soup.find_all('a', href=True):
    print(link)
    if 'xlsx' in link['href']:
#        print(link['href'])
        url="https://www.elections.on.ca/en/resource-centre/elections-results.html#accordion2022ge"+link['href']
#    print(url)
        file= url.split("/")[-1].split(".")[0]+".xlsx"
        print(file)
        urllib.request.urlretrieve(url, file) 

谢谢!

【问题讨论】:

    标签: python excel beautifulsoup


    【解决方案1】:

    修复它。请看下面的代码:

    for link in soup.find_all('a', href=True):
    #    print(link)
        if 'xlsx' in link['href']:
            print(link['href'])
            url="https://www.elections.on.ca/"+link['href']
    #        print(url)
            file= url.split("/")[-1].split(".")[0]+".xlsx"
    #        print(file)
            urllib.request.urlretrieve(url, file)  
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2021-04-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-09-29
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多