【问题标题】:How to skip scraping with same element with Beautifulsoup4如何使用 Beautifulsoup4 跳过使用相同元素的抓取
【发布时间】:2019-05-20 04:19:34
【问题描述】:

我想从网页中抓取视频,但该网页中有两个 iframe 标签。 一个用于显示 Facebook 页面,另一个用于嵌入视频。 我只想从中获取视频网址.. 但是当我尝试抓取时,我得到了所有 iframe..

像这样:

url_videos = requests.get(link_to_video)

video_link = BeautifulSoup(url_videos.text, 'html.parser')

video_on_iframe = video_link.find('iframe')

print(video_on_iframe)

当我尝试运行上面的代码时,我得到了这个结果:

<iframe allow="encrypted-media" allowtransparency="true" frameborder="0" height="80" scrolling="no" src="https://www.facebook.com/plugins/page.php?href=https%3A%2F%2Fwww.facebook.com%2FAnimeindoFans%2F&amp;tabs&amp;width=280&amp;height=180&amp;small_header=true&amp;adapt_container_width=true&amp;hide_cover=true&amp;show_facepile=false&amp;appId=123434497681677" style="border:none;overflow:hidden" width="280"></iframe>
<iframe allow="encrypted-media" allowtransparency="true" frameborder="0" height="80" scrolling="no" src="https://www.facebook.com/plugins/page.php?href=https%3A%2F%2Fwww.facebook.com%2FAnimeindoFans%2F&amp;tabs&amp;width=280&amp;height=180&amp;small_header=true&amp;adapt_container_width=true&amp;hide_cover=true&amp;show_facepile=false&amp;appId=123434497681677" style="border:none;overflow:hidden" width="280"></iframe>
<iframe allow="encrypted-media" allowtransparency="true" frameborder="0" height="80" scrolling="no" src="https://www.facebook.com/plugins/page.php?href=https%3A%2F%2Fwww.facebook.com%2FAnimeindoFans%2F&amp;tabs&amp;width=280&amp;height=180&amp;small_header=true&amp;adapt_container_width=true&amp;hide_cover=true&amp;show_facepile=false&amp;appId=123434497681677" style="border:none;overflow:hidden" width="280"></iframe>
<iframe frameborder="0" height="380" scrolling="no" src="http://www.mp4upload.com/embed-q7xxgge1yu1c.html" type="text/html" width="640">
</iframe>
<iframe allow="encrypted-media" allowtransparency="true" frameborder="0" height="80" scrolling="no" src="https://www.facebook.com/plugins/page.php?href=https%3A%2F%2Fwww.facebook.com%2FAnimeindoFans%2F&amp;tabs&amp;width=280&amp;height=180&amp;small_header=true&amp;adapt_container_width=true&amp;hide_cover=true&amp;show_facepile=false&amp;appId=123434497681677" style="border:none;overflow:hidden" width="280"></iframe>
<iframe allow="encrypted-media" allowtransparency="true" frameborder="0" height="80" scrolling="no" src="https://www.facebook.com/plugins/page.php?href=https%3A%2F%2Fwww.facebook.com%2FAnimeindoFans%2F&amp;tabs&amp;width=280&amp;height=180&amp;small_header=true&amp;adapt_container_width=true&amp;hide_cover=true&amp;show_facepile=false&amp;appId=123434497681677" style="border:none;overflow:hidden" width="280"></iframe>

我不需要那个 Facebook iframe,我只需要来自其他 iframe 的视频 URL,其属性为 height="380" 和 width="280"

当我尝试在 find() 方法中指定更多细节时:

video_on_iframe = video_link.find('iframe', width=640, height=380)

我知道了:

None
None
None
<iframe frameborder="0" height="380" scrolling="no" src="http://www.mp4upload.com/embed-q7xxgge1yu1c.html" type="text/html" width="640">
</iframe>
None
None

一个 iframe 元素,其他元素都没有..

所以..我的问题是如何找到所有iframe', width=640, height=380 值并跳过其他None 结果..?

【问题讨论】:

    标签: html python-3.x iframe web-scraping beautifulsoup


    【解决方案1】:

    您还可以要求存在src 属性:

    video_on_iframe = video_link.find('iframe', src=True)
    

    或者,结合对width 和height 的检查:

    video_on_iframe = video_link.find('iframe', src=True, width=640, height=380)
    

    【讨论】:

    • Hy.. @alecxe,感谢您对我的问题的回复,这是您第二次回答我的问题。但是在这一次中,我仍然得到了与我的代码相同的结果在上面,并没有效果。
    【解决方案2】:

    您可以使用find_all 查找所有具有该尺寸并具有 src 属性的视频。

    video_on_iframe = [video["src"] for video in video_link.find_all('iframe', width=640, 
    height=380, src=True)]
    print(video_on_iframe)
    

    [u'http://www.mp4upload.com/embed-q7xxgge1yu1c.html'] [0.2s完成]

    【讨论】:

    • 我得到的不是 None,而是带有该代码 @Zroq 的空列表
    • 现在我用这个代码:video_on_iframe = video_link.find('iframe', allow='encrypted-media' == False) if video_on_iframe is not None: URL_VIDEOS = video_on_iframe['src'] 但是这个过程很慢..
    【解决方案3】:
    video_on_frame = video_link.find_all('iframe', height = '380')## This means I wanna scrape iframe who has height value 380 . You can also use widht. 
    link_array = []
    for link in video_on_frame:  ## Your html has 1 iframe in video_on_frame format.
    
            get_iframe_url = link['src'] ## find iframe's src 
               
    
    
            try:
                link_array.append(get_iframe_url) ## add src into a array
    
            except:
                 link_array.append('Error')
    

    print(link_array) 会显示你想要的 url

    【讨论】:

    • 嘿@Omer,感谢您对我的问题的关注,但它仍然不起作用 Omer,而不是 None 我得到了空 list.
    • 现在我用这个代码:video_on_iframe = video_link.find('iframe', allow='encrypted-media' == False) if video_on_iframe is not None: URL_VIDEOS = video_on_iframe['src'] 但是这个过程很慢..
    猜你喜欢
    • 2018-02-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-11-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-07-20
    相关资源
    最近更新 更多