【问题标题】:Python - Get a list of URLs from a complicated html file for scraping purposesPython - 从复杂的 html 文件中获取 URL 列表以进行抓取
【发布时间】:2022-01-19 02:55:09
【问题描述】:

我是网络抓取的新手,无法从该网站获取“a”标签中的 URL 列表:http://www.tauntondevelopment.org//msip/JHRindex.htm。我得到的只是一个空列表-客户列表:[] 感谢您的帮助!

这是我的代码:

from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup

# This is the url of one major industrial park that we will be scraping
park_url = "http://www.tauntondevelopment.org//msip/JHRindex.htm"

uPark = uReq(park_url)
park_html = uPark.read()
uPark.close()

park_soup = soup(park_html, "html.parser")

filename = "ParkText.html"
f = open(filename, "w") 
f.write(park_soup.prettify())
f.close()

# get a list of the urls of park_url    
clients_list = []
for link in park_soup.findAll('li'):
    clients_list.append(link.get('href'))

print("clients list:", clients_list)

# write clients to a file 
filename = "taunton_JHR.csv"

f = open(filename, "w") # 
headers = "Name, Email, Address\n"
f.write(headers)
 
for client_url in clients_list:
    # call the function to scrape the individual park data
    client_url = "http://www.tauntondevelopment.org/msip/" + client_url
    try: 
        uClient = uReq(client_url)
    except:
        print("Error: Unable to open url")
        continue # continue to the next client_url in the list

    client_name, client_email, client_address = scrapeIndPark(uClient)
    
    f.write(client_name  + "," + client_email + "," + client_address + "\n")
    
f.close()


【问题讨论】:

    标签: python web-scraping beautifulsoup


    【解决方案1】:

    您是否尝试查看您下载的 html?

    <html>
     <head>
      <title>
       John Hancock Road
      </title>
      <meta content="text/html; charset=utf-8" http-equiv="Content-Type"/>
     </head>
     <frameset bordercolor="#E0E0E0" cols="25%,591*">
      <frame name="index" src="JHRleft.htm" target="content"/>
      <frame name="content" src="JHRright.htm"/>
     </frameset>
     <noframes>
      <body bgcolor="#FFFFFF">
      </body>
     </noframes>
     <frameset>
     </frameset>
    </html>
    

    请注意(至少在我的情况下)它是空的!这是因为页面是用框架构建的。要访问框架,您需要转到页面,运行网络检查器,转到 network 选项卡并查看发送后一个请求(用数据填充框架)的 url。在这种情况下,您搜索的网址可能是 http://www.tauntondevelopment.org//msip/JHRleft.htm

    【讨论】:

    • 你是对的!这是我的问题的根本原因。感谢您提供有关框架概念的信息:-)。使用您的两个答案 (kosciej16) 和 João A. Veiga 的答案,我能够生成列表。谢谢你们俩。
    【解决方案2】:

    在您的代码中,您尝试从 li 元素本身获取 href 属性。实际上,li 元素有一个嵌套的 p 和一个嵌套的 b,其中有一个嵌套的内部,你需要得到那个嵌套的 a。

    这是一个建议:

    clients_list = []
    for link in park_soup.findAll('li'):
        href_attr = link.findAll('p')[0].findAll('b')[0].findAll('a').get('href')
        clients_list.append(href_attr)
    

    另一个想法是直接获取所有标签:

    clients_list = []
    for link in park_soup.findAll('a'):
        href_attr = link.get('href')
        clients_list.append(href_attr)
    

    【讨论】:

    • 其实我觉得这不是根本情况,看我的回答
    • 感谢您的反馈。但是,这两个脚本都返回一个空列表!
    猜你喜欢
    • 1970-01-01
    • 2020-07-28
    • 2022-01-05
    • 1970-01-01
    • 1970-01-01
    • 2019-07-16
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多