【发布时间】:2022-01-19 02:55:09
【问题描述】:
我是网络抓取的新手,无法从该网站获取“a”标签中的 URL 列表:http://www.tauntondevelopment.org//msip/JHRindex.htm。我得到的只是一个空列表-客户列表:[] 感谢您的帮助!
这是我的代码:
from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup
# This is the url of one major industrial park that we will be scraping
park_url = "http://www.tauntondevelopment.org//msip/JHRindex.htm"
uPark = uReq(park_url)
park_html = uPark.read()
uPark.close()
park_soup = soup(park_html, "html.parser")
filename = "ParkText.html"
f = open(filename, "w")
f.write(park_soup.prettify())
f.close()
# get a list of the urls of park_url
clients_list = []
for link in park_soup.findAll('li'):
clients_list.append(link.get('href'))
print("clients list:", clients_list)
# write clients to a file
filename = "taunton_JHR.csv"
f = open(filename, "w") #
headers = "Name, Email, Address\n"
f.write(headers)
for client_url in clients_list:
# call the function to scrape the individual park data
client_url = "http://www.tauntondevelopment.org/msip/" + client_url
try:
uClient = uReq(client_url)
except:
print("Error: Unable to open url")
continue # continue to the next client_url in the list
client_name, client_email, client_address = scrapeIndPark(uClient)
f.write(client_name + "," + client_email + "," + client_address + "\n")
f.close()
【问题讨论】:
标签: python web-scraping beautifulsoup