【问题标题】:Issue with web scraping Job search sites网络抓取问题 求职网站
【发布时间】:2021-01-21 00:56:43
【问题描述】:

我正在网上抓取 Monster 工作网站,搜索目标为“软件开发人员”,我的目标是仅打印出 Python 终端中描述中列出了“python”的工作,同时丢弃所有Java、HTML、CSS 等的其他作业。但是,当我运行此代码时,我最终会打印页面上的所有作业。 为了解决这个问题,我创建了一个变量(称为“搜索”),它使用“python”搜索所有作业并将其转换为小写。我还创建了一个变量(称为“python_jobs”),其中包含页面上的所有工作列表。

然后我创建了一个“for”循环来查找在“python_jobs”中找到“search”的每个实例。但是,这给出了与以前相同的结果,并且无论如何都会打印出页面上的每个作业列表。有什么建议吗?

import requests
from bs4 import BeautifulSoup

URL = "https://www.monster.com/jobs/search/?q=Software-Developer"
page = requests.get(URL)
print(page)

soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(id="ResultsContainer")

search = results.find_all("h2", string=lambda text: "python" in text.lower())
python_jobs = results.find_all("section", class_="card-content")

print(len(search))

for search in python_jobs:
    title = search.find("h2", class_="title")
    company = search.find("div", class_="company")
    if None in (title, company):
        continue
    print(title.text.strip())
    print(company.text.strip())
    print()

【问题讨论】:

  • 首先你可以使用print()来查看变量中的值以及代码的哪一部分被执行——它被称为"print debuging"——也许title, company不是None所以它不会跳过它
  • 您将带有python 的元素分配给变量search,但您从不使用它——您只使用python_jobs 并将其分配给同一个变量search,因此它会从@ 中删除以前的内容987654330@。也许你应该使用for item in search 而不是for search in python_jobs。或者也许你应该以不同的方式工作 - 首先在results 中找到所有具有h2 with python` 的元素,然后使用这些元素(而不是reasults)来查找`sections 和其他信息。

标签: python web-scraping


【解决方案1】:

您的问题是您有两个不相关的单独列表 search 和 python_jobs。后来你甚至不使用列表search。您应该从 python_jobs 获取每个项目并在该项目中搜索 python。

import requests
from bs4 import BeautifulSoup

URL = "https://www.monster.com/jobs/search/?q=Software-Developer"
page = requests.get(URL)
print(page)

soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(id="ResultsContainer")

all_jobs = results.find_all("section", class_="card-content")

for job in all_jobs:
    python = job.find("h2", string=lambda text: "python" in text.lower())
    if python:
        title = job.find("h2", class_="title")
        company = job.find("div", class_="company")
        print(title.text.strip())
        print(company.text.strip())
        print()

或

import requests
from bs4 import BeautifulSoup

URL = "https://www.monster.com/jobs/search/?q=Software-Developer"
page = requests.get(URL)
print(page)

soup = BeautifulSoup(page.content, "html.parser")
results = soup.find(id="ResultsContainer")

all_jobs = results.find_all("section", class_="card-content")

for job in all_jobs:
    title = job.find("h2")
    if title:
        title = title.text.strip()
        if 'python' in title.lower():
            company = job.find("div", class_="company").text.strip()
            print(title)
            print(company)
            print()

【讨论】:

    猜你喜欢
    • 2020-05-30
    • 1970-01-01
    • 2023-03-17
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-10-21
    相关资源
    最近更新 更多