【问题标题】:Extracting p from div class python to get addresses从 div 类 python 中提取 p 以获取地址
【发布时间】:2020-09-10 19:52:18
【问题描述】:

当前代码:查找所有健身房的 url 并放入 csv 中,如下所示:

https://www.lifetime.life/life-time-locations/al-vestavia-hills.html
https://www.lifetime.life/life-time-locations/az-biltmore.html

我想要它做什么:我无法从每个 url 中提取地址。我在地址部分的尝试位于下面“代码”底部的第 4 行和第 5 行。确切的错误是:

gymrow.append(address_line1[0].text)
IndexError: list index out of range

代码*:

import urllib2
import BeautifulSoup

initial_url = "https://www.lifetime.life"

request = urllib2.Request("https://www.lifetime.life/view-all-locations.html")
response = urllib2.urlopen(request)
soup = BeautifulSoup.BeautifulSoup(response)
with open('gyms2.csv', 'w') as gf:
  gymwriter = csv.writer(gf)
  for a in soup.findAll('a'):
    if '/life-time-locations/' in a['href']:
      gymurl1 = (urlparse.urljoin(initial_url, a.get('href')))
      sitemap_content = requests.get(gymurl1).content
      gymrow = [gymurl1]

      address_line1 = soup.select('p[class~=small m-b-sm p-t-1] > span[class~=btn-icon-text]')
      gymrow.append(address_line1[0].text)

      print(gymrow)
      gymwriter.writerow(gymrow)
      time.sleep(3)

Image of inspect element: the p class, span class and the address I want to scrape

非常感谢!

【问题讨论】:

  • 首先检查您得到的响应。服务器可能会为不同的设备(手机、平板电脑、台式机)发送不同的 HTML,您可能需要 User-Agent 标头。或者它可能会阻止您的代码,并且可能会以 HTML 格式发送警告消息 - 所以将 response HTML 保存在文件中并在浏览器中打开以查看您得到的结果。
  • 我不明白你为什么使用urllib2.Requests 和requests.get 如果可以使用其中之一。
  • 您从子页面读取数据,但您没有创建soup,而是在主页上搜索。

标签: python html python-2.7 web-scraping beautifulsoup


【解决方案1】:

您从子页面获取 HTML,但未转换为 soup,因此您在主页上搜索

response = requests.get(gymurl)
sub_soup = BeautifulSoup(response.text)

我也遇到了 CSS 选择器的问题

address_line = sub_soup.select('p.small.m-b-sm.p-t-1 span.btn-icon-text')

有些页面在这个地方没有元素,它会引发错误,所以我使用try/except 来捕捉它。


在 Python 3 上测试,因为在 Python 2 上 .select() 对我不起作用

import requests
from bs4 import BeautifulSoup
import urllib.parse
import csv
import time

initial_url = "https://www.lifetime.life"

response = requests.get("https://www.lifetime.life/view-all-locations.html")
soup = BeautifulSoup(response.text)

with open('gyms2.csv', 'w') as gf:
    gymwriter = csv.writer(gf)
    for a in soup.findAll('a'):
        if '/life-time-locations/' in a['href']:
            gymurl = urllib.parse.urljoin(initial_url, a.get('href'))
            print(gymurl)

            response = requests.get(gymurl)
            sub_soup = BeautifulSoup(response.text)

            try:
                address_line = sub_soup.select('p.small.m-b-sm.p-t-1 span.btn-icon-text')
                gymrow = [gymurl, address_line[0].text.strip()]
                print(gymrow)
                gymwriter.writerow(gymrow)
                time.sleep(3)
            except Exception as ex:
                print(ex)

编辑: Python 2 使用 find() 而不是 select()

import requests
import BeautifulSoup
import csv
import urllib2
import time

initial_url = "https://www.lifetime.life"

response = requests.get("https://www.lifetime.life/view-all-locations.html")
soup = BeautifulSoup.BeautifulSoup(response.text)

with open('gyms2.csv', 'w') as gf:
    gymwriter = csv.writer(gf)
    for a in soup.findAll('a'):
        if '/life-time-locations/' in a['href']:
            gymurl = urllib2.urlparse.urljoin(initial_url, a.get('href'))
            print(gymurl)

            response = requests.get(gymurl)
            sub_soup = BeautifulSoup.BeautifulSoup(response.text)

            try:
                address_line = sub_soup.find('p', {'class': 'small m-b-sm p-t-1'}).find('span', {'class': 'btn-icon-text'})
                gymrow = [gymurl, address_line.text]
                print(gymrow)
                gymwriter.writerow(gymrow)
                time.sleep(3)
            except Exception as ex:
                print(ex)

编辑: 页面似乎有很多版本。每个页面可能需要分隔try/except。但是,如果第一个 try 工作正常,我会使用 continue 跳过下一个 try/except,而不是将第二个 try/except 放在第一个 except 中。

import requests
from bs4 import BeautifulSoup
import urllib.parse
import csv
import time

initial_url = "https://www.lifetime.life"

response = requests.get("https://www.lifetime.life/view-all-locations.html")
soup = BeautifulSoup(response.text)

with open('gyms2.csv', 'w') as gf:
    gymwriter = csv.writer(gf)
    for a in soup.findAll('a'):
        if '/life-time-locations/' in a['href']:
            gymurl = urllib.parse.urljoin(initial_url, a.get('href'))
            print(gymurl)

            response = requests.get(gymurl)
            sub_soup = BeautifulSoup(response.text)

            try:
                address_line = sub_soup.select('p.small.m-b-sm.p-t-1 span.btn-icon-text')
                gymrow = [gymurl, address_line[0].text.strip()]
                print('type 1:', gymrow)
                gymwriter.writerow(gymrow)
                time.sleep(3)
                continue # go back to `for`            
            except Exception as ex:
                print('ex:', ex)

            try:
                address_line = sub_soup.find('div', {'class': 'btn-resp-md'}).find('p')
                gymrow = [gymurl, address_line.text.strip()]
                print('type 2:', gymrow)
                gymwriter.writerow(gymrow)
                time.sleep(3)
                continue # go back to `for`            
            except Exception as ex:
                print('ex:', ex)

            try:
                address_line = sub_soup.find('p', {'class': 'm-b-grid'})
                gymrow = [gymurl, address_line.text.strip()]
                print('type 3:', gymrow)
                gymwriter.writerow(gymrow)
                time.sleep(3)
                continue # go back to `for`            
            except Exception as ex:
                print('ex:', ex)

【讨论】:

猜你喜欢
  • 1970-01-01
  • 2021-10-10
  • 2021-04-03
  • 1970-01-01
  • 2018-07-14
  • 2014-03-02
  • 1970-01-01
  • 2020-01-31
  • 2017-08-31
相关资源
最近更新 更多