使用 BeautifulSoup 进行抓取移动到下一页答案

【问题标题】：Moving to the next page using BeautifulSoup for scraping使用 BeautifulSoup 进行抓取移动到下一页
【发布时间】：2021-02-12 11:07:24
【问题描述】：

我需要从网站上抓取内容（只是标题）。我为一页做了，但我需要为网站上的所有页面都做。目前，我正在执行以下操作：

import bs4, requests
import pandas as pd
import re

headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/61.0.3163.100 Safari/537.36'}
    
    
r = requests.get(website, headers=headers)
soup = bs4.BeautifulSoup(r.text, 'html')


title=soup.find_all('h2')

我知道，当我转到下一页时，url 会发生如下变化：

website/page/2/
website/page/3/
... 
website/page/49/
...

我尝试使用 next_page_url = base_url + next_page_partial 构建递归函数，但它没有移动到下一页。

if soup.find("span", text=re.compile("Next")):
    page = "https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina/page/".format(page_num)
    page_num +=10 # I should scrape all the pages so maybe this number should be changed as I do not know at the beginning how many pages there are for that section
    print(page_num)
else:
    break

我关注了这个问题（和答案）：Moving to next page for scraping using BeautifulSoup

如果您需要更多信息，请告诉我。非常感谢

更新代码：

import bs4, requests
import pandas as pd
import re

headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/61.0.3163.100 Safari/537.36'}

page_num=1
website="https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina"

while True:
  r = requests.get(website, headers=headers)
  soup = bs4.BeautifulSoup(r.text, 'html')


  title=soup.find_all('h2')

  if soup.find("span", text=re.compile("Next")):
      page = f"https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina/page/{page_num}".format(page_num)
      page_num +=10
  else:
      break

【问题讨论】：

我相信 .format 行需要一个 {} 在你希望 page_num 去的字符串中。这有什么帮助吗？
page_num之前好像加了“page/”。
使用.format时，.format()中不需要加{}。
没有朋友，如"website/page/{}".format(page_num) 我认为虽然我个人使用 f 字符串所以它会是f"website/page/{page_num}"
哦抱歉，太习惯f字符串了

标签： python web-scraping beautifulsoup web-crawler

【解决方案1】：

如果您使用f"url/{page_num}"，则删除format(page_num)。

你可以在下面使用任何你想要的东西：

page = f"https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina/page/{page_num}"

或

page = "https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina/page/{}".format(page_num)

祝你好运！

最终答案是这样的：

import bs4, requests
import pandas as pd
import re

headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/61.0.3163.100 Safari/537.36'}

page_num=1
website="https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina"

while True:
  r = requests.get(website, headers=headers)
  soup = bs4.BeautifulSoup(r.text, 'html')


  title=soup.find_all('h2')

  if soup.find("span", text=re.compile("Next")):
      website = f"https://catania.liveuniversity.it/notizie-catania-cronaca/cronacacatenesesicilina/page/{page_num}"
      page_num +=1
  else:
      break

【讨论】：

哦，对不起。我没有检查那个。您将 url 更新为 page，但您在请求中使用了 website。请检查一下。
修改后的报废url保存在page，但是当while再次启动时，r=requests()仍然从website获取数据。
请尝试在if 条件下将page 更改为website。你还有.format(page_num)，当你使用f""时应该删除它。
page_num += 1 将是答案。在您的第一次迭代中，将使用初始化的 website。之后，page_num 在您的第二次迭代中将为 2，因为 1 已添加到您的初始值 1。
我更新了我的答案。请看一下，让我知道有问题。