【问题标题】:Scrape Facebook likes刮脸书喜欢
【发布时间】:2018-08-10 06:36:40
【问题描述】:

我想抓取网站的点赞数。使用 BeautifulSoup,这是我目前得到的:

user = 'LazadaMalaysia'

url = 'https://www.facebook.com/'+ user
response = requests.get(url)
soup = BeautifulSoup(response.content,'lxml')
f = soup.find('div', attrs={'class': '_4bl9'})

我收到的 f 输出如下:

<div class="_4bl9 _3bcp"><div aria-keyshortcuts="Alt+/" aria-label="Pembantu Navigasi" class="_6a _608n" id="u_0_8" role="menubar"><div class="_6a uiPopover" id="u_0_9"><a aria-expanded="false" aria-haspopup="true" class="_42ft _4jy0 _55pi _2agf _4o_4 _63xb _p _4jy3 _517h _51sy" href="#" id="u_0_a" rel="toggle" role="button" style="max-width:200px;"><span class="_55pe">Bahagian-bahagian pada halaman ini</span><span class="_4o_3 _3-99"><i class="img sp_m7lN5cdLBIi sx_d3bfaf"></i></span></a></div><div class="_6a _3bcs"></div><div class="_6a mrm uiPopover" id="u_0_b"><a aria-expanded="false" aria-haspopup="true" class="_42ft _4jy0 _55pi _2agf _4o_4 _3_s2 _63xb _p _4jy3 _4jy1 selected _51sy" href="#" id="u_0_c" rel="toggle" role="button" style="max-width:200px;" tabindex="-1"><span class="_55pe">Bantuan Kebolehcapaian</span><span class="_4o_3 _3-99"><i class="img sp_m7lN5cdLBIi sx_0a4c0e"></i></span></a></div></div></div>

我使用了此链接中的代码:How do I scrape the about section of a Facebook page?

不幸的是,它不起作用,我无法理解为什么会这样。这是我要抓取的部分:

【问题讨论】:

  • 看起来_4b19 类有多个divs,因为您使用的是find(),所以它只是为您提供页面中该类的第一个实例。我要说明的一点是,您想要的特定div 可能在页面的下方,而您还没有抓住它。使用find_all() 将为您提供所有类_4b19 的列表,尝试查看该列表以查看您是否获得了喜欢,或者您可能需要修改find() 的参数
  • 有同样的想法,不幸的是 find_all() 返回一个空列表,因此也无法正确过滤掉它。尝试使用“people like this”字符串和 re.search,因为它在页面源中是唯一的。
  • 试一试soup.find_all("div", string="people like this")?
  • 仍然返回一个空列表。 =/
  • 不要刮脸书,这是不允许的。您可以简单地使用 API:developers.facebook.com/tools/…

标签: python facebook web-scraping beautifulsoup


【解决方案1】:

我发现以下内容很容易使用。它应该获取喜欢和关注的数量。

import re
import requests
from bs4 import BeautifulSoup

def get_info(user,url):
    response = requests.get(f'{url}{user}')
    soup = BeautifulSoup(response.text,'lxml')
    like = soup.find("div",text=re.compile('people like this')).text
    follow = soup.find("div",text=re.compile('people follow this')).text
    print(f'likes: {like}\nfollows: {follow}\n')

if __name__ == '__main__':
    url = "https://www.facebook.com/"
    users = ['LazadaMalaysia','ronaldo','rihanna']
    [get_info(user,url) for user in users]

【讨论】:

    【解决方案2】:

    喜欢在类“_4-u3 _5sqi _5sqk”内的span标签中。这是提取喜欢的代码。

    import requests
    from bs4 import BeautifulSoup
    user = 'LazadaMalaysia'
    url = 'https://www.facebook.com/'+ user
    response = requests.get(url)
    soup = BeautifulSoup(response.content,'lxml')
    f = soup.find('div', attrs={'class': '_4-u3 _5sqi _5sqk'})
    likes=f.find('span',attrs={'class':'_52id _50f5 _50f7'}) #finding span tag inside class
    print(likes.text)
    

    希望我已经解决了你的问题。

    【讨论】:

    • @shivank01 你知道如何取消关注者吗?我想我只需要更改课程代码,但我不知道如何...提前谢谢您!
    猜你喜欢
    • 1970-01-01
    • 2012-10-09
    • 1970-01-01
    • 2016-06-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2012-06-28
    相关资源
    最近更新 更多