【发布时间】:2022-06-14 23:01:09
【问题描述】:
我正在尝试从代码中看到的链接下载文件。但是,当我打开下载的文件时出现以下错误。我该如何解决这个问题?
请看下面的代码:
import os
import requests
from bs4 import BeautifulSoup
# Python 3.x
from urllib.request import urlopen, urlretrieve, quote
from urllib.parse import urljoin
import urllib
headers={"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.71 Safari/537.36"}
resp = requests.get("https://www.elections.on.ca/en/resource-centre/elections-results.html#accordion2022ge")
soup = BeautifulSoup(resp.text,"html.parser")
for link in soup.find_all('a', href=True):
print(link)
if 'xlsx' in link['href']:
# print(link['href'])
url="https://www.elections.on.ca/en/resource-centre/elections-results.html#accordion2022ge"+link['href']
# print(url)
file= url.split("/")[-1].split(".")[0]+".xlsx"
print(file)
urllib.request.urlretrieve(url, file)
谢谢!
【问题讨论】:
标签: python excel beautifulsoup