【问题标题】:Get Web data with images for HTML table获取带有 HTML 表格图像的 Web 数据
【发布时间】:2022-11-05 16:55:29
【问题描述】:

我正在尝试使用来自this link 的图像提取文章正文,以便使用提取的文章正文制作 HTML 表格。所以,我尝试使用BeautifulSoup。

t_link = 'https://www.cnbc.com/2022/01/03/5-ways-to-reset-your-retirement-savings-and-save-more-money-in-2022.html'
page = requests.get(t_link)
soup_page = BeautifulSoup(page.content, 'html.parser')


html_article = soup_page.find_all("div", {"class": re.compile('ArticleBody-articleBody.?')})


for article_body in html_article: 
  print(article_body)

但不幸的是,article_body 没有显示任何图像,就像这样。因为,<div class="InlineImage-wrapper"> 不是这样刮的

那么,如何获取带有文章图片的文章数据,以便制作 HTML 表格呢?

【问题讨论】:

  • 似乎该站点使用延迟加载方法来加载图像,这意味着它是在渲染页面时加载的,我认为 bs4 无法处理,因为它不渲染页面(它只抓取源页面,而不是渲染页)
  • 图片有<div class="InlineImage-wrapper">,我是初学者,所以我遇到了问题
  • 是的,正如我告诉你的那样,图像的 HTML 标记在那里,但是图像没有在服务器端加载,它在客户端呈现(它使用延迟加载),bs4 无法直接检索图像,因为它不会渲染图像。我尝试检查页面,仍然有一种使用 bs4 的方法,但是您需要使用来自例如的 ID。 id="ArticleBody-InlineImage-106967852" = 106967852,并在window.__s_data上找到它的映射,找到映射后,从该对象中获取图像
  • 我不知道以何种方式获取图像(延迟加载,请求看不到它,因为它是从不同的源动态加载的,但是存在于ld+json 脚本标签等中 - 请参阅@baduker 的回复)将有助于 HTML 表格...?您抓取数据以对其进行处理,对其进行分析,无论如何,而不是“抓取 HTML 以创建 HTML...表”。无意冒犯,但您的问题存在严重的逻辑缺陷。
  • @BarrythePlatipus 是的,实际上,我是初学者(不是开发人员或类似的人),我正在寻找是否有办法抓取文章内容(包含所有段落和图像)。我认为几乎所有东西都可以报废,在 python 中有很多库可以做这些类型的东西,这对我来说是未知的,任何人都可以解决我的问题。我从 baduker 的回复中得到了一个想法,特别感谢他。从他的回答来看,我正试图以另一种方式解决我的问题。另外,非常感谢 Barry 的友好回复。

标签: python web-scraping beautifulsoup


【解决方案1】:

我不太明白你的目标,所以我的可能不是你想要的答案。

在该页面的 html 源代码中,您可以在底部的脚本中找到所有内容。

它具有 JSON 格式的页面内容。 如果您只是使用 grep 和 jq(一个很棒的 JSON cli 实用程序),您可以运行

curl -kL "https://www.cnbc.com/2022/01/03/5-ways-to-reset-your-retirement-savings-and-save-more-money-in-2022.html" | 
grep -Po '"body":.+"body".' | 
grep -Po '{"content":[.+"body".' | 
jq '[.content[]|select(.tagName|contains("image"))]'

获取有关图像的所有信息

[
  {
    "tagName": "image",
    "attributes": {
      "id": "106967852",
      "type": "image",
      "creatorOverwrite": "PM Images",
      "headline": "Retirement Savings",
      "url": "https://image.cnbcfm.com/api/v1/image/106967852-1635524865061-GettyImages-1072593728.jpg?v=1635525026",
      "datePublished": "2021-10-29T16:30:26+0000",
      "copyrightHolder": "PM Images",
      "width": "2233",
      "height": "1343"
    },
    "data": {
      "__typename": "image"
    },
    "children": [],
    "__typename": "bodyContent"
  },
  {
    "tagName": "image",
    "attributes": {
      "id": "106323101",
      "type": "image",
      "creatorOverwrite": "JGI/Jamie Grill",
      "headline": "GP: 401k money jar on desk of businesswoman",
      "url": "https://image.cnbcfm.com/api/v1/image/106323101-1578344280328gettyimages-672157227.jpeg?v=1641216437",
      "datePublished": "2020-01-06T20:58:19+0000",
      "copyrightHolder": "JGI/Jamie Grill",
      "width": "5120",
      "height": "3418"
    },
    "data": {
      "__typename": "image"
    },
    "children": [],
    "__typename": "bodyContent"
  }
]

如果您只需要 URL,请运行

curl -kL "https://www.cnbc.com/2022/01/03/5-ways-to-reset-your-retirement-savings-and-save-more-money-in-2022.html" | 
grep -Po '"body":.+"body".' | 
grep -Po '{"content":[.+"body".' | 
jq  -r '[.content[]|select(.tagName|contains("image"))]|.[].attributes.url'

要得到

https://image.cnbcfm.com/api/v1/image/106967852-1635524865061-GettyImages-1072593728.jpg?v=1635525026
https://image.cnbcfm.com/api/v1/image/106323101-1578344280328gettyimages-672157227.jpeg?v=1641216437

【讨论】:

  • 感谢您的回答,您的回答有助于提取图像。我只是想复制所有内容元素并将它们粘贴到 HTML 编辑器中以重新生成博客内容。
【解决方案2】:

您想要的一切都在源代码HTML 中,但您需要跳过几个环节才能获得该数据。

我提供以下内容:

  • 文章正文
  • 两 (2) 张图片与文章正文和标题视频 (1) 的 URL

就是这样:

import json
import re

import requests
from bs4 import BeautifulSoup

headers = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:104.0) Gecko/20100101 Firefox/104.0",
}

with requests.Session() as s:
    s.headers.update(headers)
    url = "https://www.cnbc.com/2022/01/03/5-ways-to-reset-your-retirement-savings-and-save-more-money-in-2022.html"
    script = [
        s.text for s in
        BeautifulSoup(s.get(url).text, "lxml").find_all("script")
        if "window.__s_data" in s.text
    ][0]
    payload = json.loads(
        re.match(r"window.__s_data=(.*);swindow.__c_data=", script).group(1)
    )
    article_data = (
        payload
        ["page"]
        ["page"]
        ["layout"][3]
        ["columns"][0]
        ["modules"][2]
        ["data"]
    )
    print(article_data["articleBodyText"])
    for item in article_data["body"]["content"]:
        if "url" in item["attributes"].keys():
            print(item["attributes"]["url"])

这应该打印:

  1. 整个文章正文 (为简洁起见已编辑)
    The new year offers opportunities for many Americans in their careers and financial lives. The "Great Reshuffle" is expected to continue as employees leave jobs and take new ones at a rapid clip. At the same time, many workers have made a vow to save more this year, yet many admit they don't know how they'll stick to that goal. One piece of advice: Keep it simple. 
    [...]
    

    上面提到的资产网址:

    https://www.cnbc.com/video/2022/01/03/how-to-choose-the-best-retirement-strategy-for-2022.html
    https://image.cnbcfm.com/api/v1/image/106967852-1635524865061-GettyImages-1072593728.jpg?v=1635525026
    https://image.cnbcfm.com/api/v1/image/106323101-1578344280328gettyimages-672157227.jpeg?v=1641216437
    

    编辑:

    如果要下载图像,请使用以下命令:

    import json
    import os
    import re
    from pathlib import Path
    from shutil import copyfileobj
    
    import requests
    from bs4 import BeautifulSoup
    
    headers = {
        "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:104.0) Gecko/20100101 Firefox/104.0",
    }
    
    url = "https://www.cnbc.com/2022/01/03/5-ways-to-reset-your-retirement-savings-and-save-more-money-in-2022.html"
    
    
    def download_images(image_source: str, directory: str) -> None:
        """Download images from a given source and save them to a given directory."""
        os.makedirs(directory, exist_ok=True)
        save_dir = Path(directory)
        if re.match(r".*.jp[e-g]", image_source):
            file_name = save_dir / image_source.split("/")[-1].split("?")[0]
            with s.get(image_source, stream=True) as img, open(file_name, "wb") as output:
                copyfileobj(img.raw, output)
    
    
    with requests.Session() as s:
        s.headers.update(headers)
        script = [
            s.text for s in
            BeautifulSoup(s.get(url).text, "lxml").find_all("script")
            if "window.__s_data" in s.text
        ][0]
        payload = json.loads(
            re.match(r"window.__s_data=(.*);swindow.__c_data=", script).group(1)
        )
        article_data = (
            payload
            ["page"]
            ["page"]
            ["layout"][3]
            ["columns"][0]
            ["modules"][2]
            ["data"]
        )
        print(article_data["articleBodyText"])
        for item in article_data["body"]["content"]:
            if "url" in item["attributes"].keys():
                url = item["attributes"]["url"]
                print(url)
                download_images(url, "images")
    

【讨论】:

    猜你喜欢
    • 2023-03-17
    • 2015-11-14
    • 2012-12-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-12-28
    相关资源
    最近更新 更多