【问题标题】:Scraping through all pages抓取所有页面
【发布时间】:2021-01-20 02:33:13
【问题描述】:

我正在尝试抓取以下网站:voxnews.info

import requests
from bs4 import BeautifulSoup
from concurrent.futures import ThreadPoolExecutor
import pandas as pd

web='https://voxnews.info'
def main(req, num, web):
    r = req.get(web+"/page/{}/".format(num))
    soup = BeautifulSoup(r.content, 'html.parser')
    goal = [(x.time.text, x.h1.a.get_text(strip=True), x.select_one("span.cat-links").get_text(strip=True), x.p.get_text(strip=True))
           for x in soup.select("div.site-content")]

    return goal


with ThreadPoolExecutor(max_workers=30) as executor:
    with requests.Session() as req:
        fs = [executor.submit(main, req, num) for num in range(1, 2)] # need to scrape all the webpages in the website
        allin = []
        for f in fs:
            allin.extend(f.result())
        df = pd.DataFrame.from_records(
            allin, columns=["Date", "Title", "Category", "Content"])
        print(df)

但是代码有两个问题:

  • 第一个是我没有抓取所有页面(我目前将 1 和 2 放在范围内,但我需要所有页面);
  • 它没有正确保存日期。

如果可以看看代码并告诉我如何改进它以解决这两个问题,那就太棒了。

【问题讨论】:

  • 我假设您想从页面上的所有文章中抓取数据?
  • 你好守门员 1998,是的,没错。我需要在所有页面中抓取每篇文章的日期标题类别和内容。
  • 好的,我会处理一些事情。我首先要说的是,这些文章没有我可以看到的与之关联的 p 标签。您希望日期采用什么格式?
  • 感谢守门员 1998。我已尽力写下一些可行的代码,但不幸的是,它并没有达到我的预期。感谢您的宝贵时间和帮助

标签: python web-scraping beautifulsoup


【解决方案1】:

一些小改动。

首先,没有必要对单个请求使用 requests.Session() - 您不会尝试在请求之间保存数据。

对您的with 语句的方式进行了细微更改,我不知道它是否更正确,或者我是如何做到的,您不需要在执行器仍然打开的情况下运行所有​​代码。

我为您提供了两种解析日期的选项,可以是网站上的日期、意大利语字符串,也可以是日期时间对象。

我没有在文章中看到任何“p”标签,所以我删除了那部分。似乎为了获得文章的“内容”,您必须实际导航并单独抓取它们。我从代码中删除了该行。

在您的原始代码中,您并没有获得页面上的每篇文章,而是每篇文章的第一篇。每页只有一个“div.site-content”标签,但有多个“article”标签。这就是这种变化。

最后,我更喜欢查找而不是选择,但这只是我的风格选择。这在前三页对我有用,我没有尝试整个网站。运行时要小心,30 个请求的 78 个块可能会阻止您...

import requests
from bs4 import BeautifulSoup
from concurrent.futures import ThreadPoolExecutor
import pandas as pd
import datetime


def main(num, web):
    r = requests.get(web+"/page/{}/".format(num))
    soup = BeautifulSoup(r.content, 'html.parser')
    html = soup.find("div", class_="site-content")
    articles = html.find_all("article")
    
    # Date as string In italian
    goal = [(x.time.get_text(), x.h1.a.get_text(strip=True), x.find("span", class_="cat-links").get_text(strip=True)) for x in articles]
    # OR as datetime object
    goal = [(datetime.datetime.strptime(x.time["datetime"], "%Y-%m-%dT%H:%M:%S%z"), x.h1.a.get_text(strip=True), x.find("span", class_="cat-links").get_text(strip=True)) for x in articles]

    return goal


web='https://voxnews.info'

r = requests.get(web)
soup = BeautifulSoup(r.text, "html.parser")
last_page = soup.find_all("a", class_="page-numbers")[1].get_text()
last_int = int(last_page.replace(".",""))

### BE CAREFUL HERE WITH TESTING, DON'T USE ALL 2,320 PAGES ###
with ThreadPoolExecutor(max_workers=30) as executor:
    fs = [executor.submit(main, num, web) for num in range(1, last_int)]

allin = []
for f in fs:
    allin.extend(f.result())
df = pd.DataFrame.from_records(
    allin, columns=["Date", "Title", "Category"])
print(df)

【讨论】:

  • 您必须获取每个页面的 href,导航到该页面,然后分别抓取这些页面。
猜你喜欢
  • 1970-01-01
  • 2012-01-12
  • 2014-01-10
  • 2021-09-20
  • 1970-01-01
  • 1970-01-01
  • 2021-02-05
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多