【发布时间】:2021-01-21 00:56:43
【问题描述】:
我正在网上抓取 Monster 工作网站,搜索目标为“软件开发人员”,我的目标是仅打印出 Python 终端中描述中列出了“python”的工作,同时丢弃所有Java、HTML、CSS 等的其他作业。但是,当我运行此代码时,我最终会打印页面上的所有作业。 为了解决这个问题,我创建了一个变量(称为“搜索”),它使用“python”搜索所有作业并将其转换为小写。我还创建了一个变量(称为“python_jobs”),其中包含页面上的所有工作列表。
然后我创建了一个“for”循环来查找在“python_jobs”中找到“search”的每个实例。但是,这给出了与以前相同的结果,并且无论如何都会打印出页面上的每个作业列表。有什么建议吗?
import requests
from bs4 import BeautifulSoup
URL = "https://www.monster.com/jobs/search/?q=Software-Developer"
page = requests.get(URL)
print(page)
soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(id="ResultsContainer")
search = results.find_all("h2", string=lambda text: "python" in text.lower())
python_jobs = results.find_all("section", class_="card-content")
print(len(search))
for search in python_jobs:
title = search.find("h2", class_="title")
company = search.find("div", class_="company")
if None in (title, company):
continue
print(title.text.strip())
print(company.text.strip())
print()
【问题讨论】:
-
首先你可以使用
print()来查看变量中的值以及代码的哪一部分被执行——它被称为"print debuging"——也许title, company不是None所以它不会跳过它 -
您将带有
python的元素分配给变量search,但您从不使用它——您只使用python_jobs并将其分配给同一个变量search,因此它会从@ 中删除以前的内容987654330@。也许你应该使用for item in search而不是for search in python_jobs。或者也许你应该以不同的方式工作 - 首先在results中找到所有具有h2 withpython` 的元素,然后使用这些元素(而不是reasults)来查找`sections 和其他信息。
标签: python web-scraping