【发布时间】:2021-11-15 18:46:05
【问题描述】:
我正在编写一个 python 脚本,该脚本通过使用 Selenium 和 BeautifulSoup 进行网络抓取,从我自己的 LinkedIn 个人资料中提取性能数据。
我能够通过 Chrome 成功访问我的个人资料并提取一些数据,但 cmets 似乎很棘手。
这是我目前所拥有的:
postComments = []
src = browser.page_source
#beautiful soup instance:
soup = BeautifulSoup(src, features="lxml")
bs4TagsComments = soup.find_all("li", attrs = {"class" : "social-details-social counts__item social-details-social-counts__comments"})
for tag in bs4TagsComments:
strtag = str(tag)
list_of_matches = re.findall('[,0-9]+',strtag)
last_string = list_of_matches.pop()
without_comma = last_string.replace(',','')
commentsCount = int(without_comma)
postComments.append(commentsCount)
print(postComments)
理论上,以上内容应该可以工作 - 但是,打印出来的只是一个空列表。 有评论计数要提取,如果没有,我至少应该得到一个 '0's 的字典。
【问题讨论】:
标签: python selenium web-scraping beautifulsoup