【问题标题】:Pull comment count from a LinkedIn post in python using Selenium & Beautifulsoup使用 Selenium 和 Beautifulsoup 从 Python 中的 LinkedIn 帖子中提取评论计数
【发布时间】:2021-11-15 18:46:05
【问题描述】:

我正在编写一个 python 脚本,该脚本通过使用 Selenium 和 BeautifulSoup 进行网络抓取,从我自己的 LinkedIn 个人资料中提取性能数据。

我能够通过 Chrome 成功访问我的个人资料并提取一些数据,但 cmets 似乎很棘手。

这是我目前所拥有的:

postComments = []

src = browser.page_source
#beautiful soup instance:
soup = BeautifulSoup(src, features="lxml")

bs4TagsComments = soup.find_all("li", attrs = {"class" : "social-details-social counts__item social-details-social-counts__comments"})
for tag in bs4TagsComments:
    strtag = str(tag)
    list_of_matches = re.findall('[,0-9]+',strtag)
    last_string = list_of_matches.pop()
    without_comma = last_string.replace(',','')
    commentsCount = int(without_comma)
    postComments.append(commentsCount)

print(postComments)

理论上,以上内容应该可以工作 - 但是,打印出来的只是一个空列表。 有评论计数要提取,如果没有,我至少应该得到一个 '0's 的字典。

【问题讨论】:

    标签: python selenium web-scraping beautifulsoup


    【解决方案1】:

    使用Regex 能够提取 cmets 的值。尝试如下。

    from selenium import webdriver
    from bs4 import BeautifulSoup
    import time
    import re
    
    driver = webdriver.Chrome(executable_path="path to chromedriver.exe")
    driver.implicitly_wait(10)
    driver.maximize_window()
    
    driver.get("https://www.linkedin.com/")
    time.sleep(30) # to manually login
    
    soup = BeautifulSoup(driver.page_source,'html5lib')
    regex = re.compile('.*social-details-social-counts__comments.*')
    comments = soup.find_all('li',{'class': regex}) # find all 'li' tags that has `social-details-social-counts__comments` in it.
    for comment in comments:
        value = comment.getText().replace('\n','').replace(' ', '') #  for text without whitespaces
        print(value)
    
    1comment
    1comment
    14comments
    5comments
    29comments
    4comments
    3comments
    ...
    

    用于根据帖子提取 cmets 计数:

    feeds = soup.find_all(code to find the feeds)
    
    for feed in feeds:
        regex = re.compile('.*social-details-social-counts__comments.*')
        try:
            comments = feed.find('li',{'class': regex}).getText().replace('\n','').replace(' ', '')
        except:
            comments = None
    

    【讨论】:

    • 谢谢!得到这个工作:) 如果帖子上没有 cmets,我将如何让它打印“0cmets”?
    • @ZamenaJaffer - 我们需要先找出feeds 或posts,然后遍历它们以检查是否有cmets。更新了相同的答案。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2020-10-30
    • 1970-01-01
    • 2021-09-15
    • 2021-11-21
    • 2021-06-30
    • 2022-01-10
    • 2015-09-20
    相关资源
    最近更新 更多