【问题标题】:Get next page URL获取下一页 URL
【发布时间】:2018-03-29 14:40:31
【问题描述】:

现在我尝试从网页中抓取所有 URL。共有5个分类,每个分类有不同的页面(每页有10篇文章)。

例如:

Categories   Pages
Banana          5
Apple          14
Cherry          7
Melon           6
Berry           2

代码:

import requests
from bs4 import BeautifulSoup
import re
from urllib.parse import urljoin


res = requests.get('http://www.abcde.com/SearchParts')
soup = BeautifulSoup(res.text,"lxml")
href = [ a["href"] for a in soup.findAll("a", {"id" : re.compile("parts_img.*")})]
b1 =[]
for url in href:
    b1.append("http://www.abcde.com"+url)
print (b1)

从主页“http://www.abcde.com/SearchParts”我可以抓取每个类别的首页 URL。 B1 是第一页的 URL 列表。

像这样:

Categories   Pages                       url
Banana          1     http://www.abcde.com/A
Apple           1     http://www.abcde.com/B
Cherry          1     http://www.abcde.com/C
Melon           1     http://www.abcde.com/E
Berry           1     http://www.abcde.com/F

然后我使用 b1 的源代码来抓取下一页的 URL。所以 b2 是第二页的 URL 列表。

代码:

b2=[]
for url in b1:
    res2 = requests.get(url).text
    soup2 = BeautifulSoup(res2,"lxml")
    url_n=soup2.find('',rel = 'next')['href']
    b2.append("http://www.abcde.com"+url_n)
print(b2)

像这样:

Categories   Pages                       url
    Banana          1     http://www.abcde.com/A/s=1&page=2
    Apple           1     http://www.abcde.com/B/s=9&page=2
    Cherry          1     http://www.abcde.com/C/s=11&page=2
    Melon           1     http://www.abcde.com/E/s=7&page=2
    Berry           1     http://www.abcde.com/F/s=5&page=2

现在当我尝试做第三个时,这是一个错误,因为 Berry 的第二页是最后一页,源代码中没有“下一页”。特别是当每个类别都有不同的页面/网址时,我应该怎么做?

整个代码(直到出错):

import requests
from bs4 import BeautifulSoup
import re
from urllib.parse import urljoin


res = requests.get('http://www.ca2-health.com/frontend/SearchParts')
soup = BeautifulSoup(res.text,"lxml")
href = [ a["href"] for a in soup.findAll("a", {"id" : re.compile("parts_img.*")})]
b1 =[]
for url in href:
    b1.append("http://www.ca2-health.com"+url)
print (b1)
print("===================================================")
b2=[]
for url in b1:
    res2 = requests.get(url).text
    soup2 = BeautifulSoup(res2,"lxml")
    url_n=soup2.find('',rel = 'next')['href']
    b2.append("http://www.ca2-health.com"+url_n)
print(b2)
print("===================================================")
b3=[]
for url in b2:
    res3 = requests.get(url).text
    soup3 = BeautifulSoup(res3,"lxml")
    url_n=soup3.find('',rel = 'next')['href']
    b3.append("http://www.ca2-health.com"+url_n)
print(b3)

在此之后,我会将 b1、b2、b3 和...作为一个列表,从那时起我将拥有该页面的所有 URL。

【问题讨论】:

    标签: python loops beautifulsoup


    【解决方案1】:

    我猜你收到的是KeyError。处理异常并继续循环。如果您收到 KeyError,请执行以下操作:

    try:
        url_n = soup3.find(rel='next')['href']
    except KeyError:
        continue
    

    或

    try:
        url_n = soup3.find(rel='next').get('href')
    except AttributeError:
        continue
    

    【讨论】:

      猜你喜欢
      • 2020-03-07
      • 2017-06-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2013-03-11
      • 1970-01-01
      • 2015-01-04
      • 2015-03-17
      相关资源
      最近更新 更多