【问题标题】:How to scrape HTML elements with no class or id specified in attribute with Beautifulsoup4如何使用 Beautifulsoup4 抓取属性中未指定类或 id 的 HTML 元素
【发布时间】:2018-12-18 06:06:19
【问题描述】:

我想从页面中提取单独的内容描述,我可以使用 attribute 中指定的 class 或 id 来完成。但是.. 如果在 html tag 中没有指定 class 或 id 属性,我不知道如何获取元素。

喜欢这个截图:

<div class="cat_box_desc">
    <h3>Status:</h3>
    on-going <br>
    <h3>Genres:</h3>

    <br>
    <h3>Description:</h3>
    <div align="justify">
        <p> Information</p>
        <p>Type: TV</p>
        <p>Episodes: Unknown</p>
        <p>Status: Currently Airing</p>
        <p>Aired: Oct 7, 2013 to ?</p>
        <p>Producers: Sunrise, TV Tokyo, Sotsu Agency</p>
        <p>Genres: Mecha</p>
        <p>Duration: 25 min. per episode</p>
        <p>Synopsis:</p>
        <p>Gundam Build Fighter adalah sebuah pertarungan simulasi Gundam. Unit Gundam dirangkai dari model plastiknya. Tokoh utamanya adalah seorang anak laki-laki yang bernama Iori Sei. Sei memiliki kemampuan merangkai Gundam yang hebat, namun dia tak
            memiliki kemampuan untuk mengendalikan gundam yang ia rangkai saat melakukan Gunpla Battle. Namun satu hari dia bertemu dengan seorang pencuri roti misterius, yang memberinya sebuah batu permata.</p>
    </div><br>
    <div style="padding-left: 560px; padding-bottom:20px;" class="spacebook">
        <div class="fb-like" data-href="http://animeindo.video/category/gundam-build-fighter/" data-width="450" data-layout="box_count" data-show-faces="false" data-send="false"></div>
    </div>
</div>

我可以在class="cat_box_desc"里面抓取数据,但是我会得到里面的所有数据,我不想要它,我想分离数据。

我不知道像上面截图那样分开数据有status、genre、description、information 和 H1 和 P 标签中的其他内容,因为上面没有指定 class 或 id。

那么如何在 Beautifulsoup4.. 中做到这一点呢?

【问题讨论】:

    标签: python web-scraping beautifulsoup


    【解决方案1】:

    BeautifulSoup 已经是一个非常好的选择,因为它是一个非常灵活的库,有很多定位元素的方法。

    对于:-separated 字段,我会将它们解析到字典中以便于访问:

    import re
    
    from bs4 import BeautifulSoup
    
    data = """
    <div class="cat_box_desc">
        <h3>Status:</h3>
        on-going <br>
        <h3>Genres:</h3>
    
        <br>
        <h3>Description:</h3>
        <div align="justify">
            <p> Information</p>
            <p>Type: TV</p>
            <p>Episodes: Unknown</p>
            <p>Status: Currently Airing</p>
            <p>Aired: Oct 7, 2013 to ?</p>
            <p>Producers: Sunrise, TV Tokyo, Sotsu Agency</p>
            <p>Genres: Mecha</p>
            <p>Duration: 25 min. per episode</p>
            <p>Synopsis:</p>
            <p>Gundam Build Fighter adalah sebuah pertarungan simulasi Gundam. Unit Gundam dirangkai dari model plastiknya. Tokoh utamanya adalah seorang anak laki-laki yang bernama Iori Sei. Sei memiliki kemampuan merangkai Gundam yang hebat, namun dia tak
                memiliki kemampuan untuk mengendalikan gundam yang ia rangkai saat melakukan Gunpla Battle. Namun satu hari dia bertemu dengan seorang pencuri roti misterius, yang memberinya sebuah batu permata.</p>
        </div><br>
        <div style="padding-left: 560px; padding-bottom:20px;" class="spacebook">
            <div class="fb-like" data-href="http://animeindo.video/category/gundam-build-fighter/" data-width="450" data-layout="box_count" data-show-faces="false" data-send="false"></div>
        </div>
    </div>"""
    
    soup = BeautifulSoup(data, "html.parser")
    
    # first locate the container with the desired fields
    description = soup.find("h3", text="Description:").find_next_sibling()
    
    # get all the ":"-separated fields into a dictionary 
    pattern = re.compile(r"\w+:\s.*?")
    
    data = dict(field.split(":") for field in description.find_all(text=pattern))
    
    print(data)
    

    打印:

    {'Type': ' TV', 'Episodes': ' Unknown', 'Status': ' Currently Airing', 'Aired': ' Oct 7, 2013 to ?', 'Producers': ' Sunrise, TV Tokyo, Sotsu Agency', 'Genres': ' Mecha', 'Duration': ' 25 min. per episode'}
    

    现在这不会捕获 Synopsis,因为它的值位于单独的 p 元素中,但您可以通过以下方式获取它:

    data["Synopsis"] = description.find("p", text="Synopsis:").find_next_sibling("p").get_text()
    

    完整的美化输出:

    {'Aired': ' Oct 7, 2013 to ?',
     'Duration': ' 25 min. per episode',
     'Episodes': ' Unknown',
     'Genres': ' Mecha',
     'Producers': ' Sunrise, TV Tokyo, Sotsu Agency',
     'Status': ' Currently Airing',
     'Synopsis': 'Gundam Build Fighter adalah sebuah pertarungan simulasi Gundam. '
                 'Unit Gundam dirangkai dari model plastiknya. Tokoh utamanya '
                 'adalah seorang anak laki-laki yang bernama Iori Sei. Sei '
                 'memiliki kemampuan merangkai Gundam yang hebat, namun dia tak\n'
                 '            memiliki kemampuan untuk mengendalikan gundam yang '
                 'ia rangkai saat melakukan Gunpla Battle. Namun satu hari dia '
                 'bertemu dengan seorang pencuri roti misterius, yang memberinya '
                 'sebuah batu permata.',
     'Type': ' TV'}
    

    我们在这里使用了一些技术,下面是库文档相应部分的文档链接。请务必查看以更好地了解这些功能:

    【讨论】:

      猜你喜欢
      • 2021-10-13
      • 2011-12-11
      • 2019-05-20
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-12-19
      • 1970-01-01
      • 2023-03-09
      相关资源
      最近更新 更多