【问题标题】:Extract emails from a web page in python [closed]从python中的网页中提取电子邮件[关闭]
【发布时间】:2020-10-30 22:12:43
【问题描述】:

我发现以下代码可以抓取网站(我认为是所有网站)以获取电子邮件

import re
import requests
import requests.exceptions
from urllib.parse import urlsplit
from collections import deque
from bs4 import BeautifulSoup

# starting url. replace google with your own url.
starting_url = 'http://www.miet.ac.in'

# a queue of urls to be crawled
unprocessed_urls = deque([starting_url])

# set of already crawled urls for email
processed_urls = set()

# a set of fetched emails
emails = set()

# process urls one by one from unprocessed_url queue until queue is empty
while len(unprocessed_urls):

    # move next url from the queue to the set of processed urls
    url = unprocessed_urls.popleft()
    processed_urls.add(url)

    # extract base url to resolve relative links
    parts = urlsplit(url)
    base_url = "{0.scheme}://{0.netloc}".format(parts)
    path = url[:url.rfind('/')+1] if '/' in parts.path else url

    # get url's content
    print("Crawling URL %s" % url)
    try:
        response = requests.get(url)
    except (requests.exceptions.MissingSchema, requests.exceptions.ConnectionError):
        # ignore pages with errors and continue with next url
        continue

    # extract all email addresses and add them into the resulting set
    # You may edit the regular expression as per your requirement
    new_emails = set(re.findall(r"[a-z0-9\.\-+_]+@[a-z0-9\.\-+_]+\.[a-z]+", response.text, re.I))
    emails.update(new_emails)
    print(emails)
    # create a beutiful soup for the html document
    soup = BeautifulSoup(response.text, 'lxml')

    # Once this document is parsed and processed, now find and process all the anchors i.e. linked urls in this document
    for anchor in soup.find_all("a"):
        # extract link url from the anchor
        link = anchor.attrs["href"] if "href" in anchor.attrs else ''
        # resolve relative links (starting with /)
        if link.startswith('/'):
            link = base_url + link
        elif not link.startswith('http'):
            link = path + link
        # add the new url to the queue if it was not in unprocessed list nor in processed list yet
        if not link in unprocessed_urls and not link in processed_urls:
            unprocessed_urls.append(link)

如何修改这样的代码以仅提取一个网页..?我只需要定位一个网页而不是整个网站。

【问题讨论】:

  • 您的问题不清楚或仅限于特定问题,请查看How to Ask 并使用确切问题编辑您的问题
  • 从网站(不是所有网站,只有一个链接)获取电子邮件的问题
  • 删除循环似乎很简单。老实说......然后更容易不打扰队列,只需拔出 bs4 位并发出单个请求。
  • @QHarr 我是 python 的新手 :)
  • @QHarr 你能帮我吗.. 该网站将有一个 JavaScript 并且代码不处理它?有解决办法吗?

标签: python beautifulsoup python-requests


【解决方案1】:

只需删除从for anchor in soup.find_all("a"): 开始的所有行。您的文档应如下所示:

import re
import requests
import requests.exceptions
from urllib.parse import urlsplit
from collections import deque
from bs4 import BeautifulSoup

# starting url. replace google with your own url.
starting_url = 'http://www.miet.ac.in'

# a queue of urls to be crawled
unprocessed_urls = deque([starting_url])

# set of already crawled urls for email
processed_urls = set()

# a set of fetched emails
emails = set()

# process urls one by one from unprocessed_url queue until queue is empty
while len(unprocessed_urls):

    # move next url from the queue to the set of processed urls
    url = unprocessed_urls.popleft()
    processed_urls.add(url)

    # extract base url to resolve relative links
    parts = urlsplit(url)
    base_url = "{0.scheme}://{0.netloc}".format(parts)
    path = url[:url.rfind('/')+1] if '/' in parts.path else url

    # get url's content
    print("Crawling URL %s" % url)
    try:
        response = requests.get(url)
    except (requests.exceptions.MissingSchema, requests.exceptions.ConnectionError):
        # ignore pages with errors and continue with next url
        continue

    # extract all email addresses and add them into the resulting set
    # You may edit the regular expression as per your requirement
    new_emails = set(re.findall(r"[a-z0-9\.\-+_]+@[a-z0-9\.\-+_]+\.[a-z]+", response.text, re.I))
    emails.update(new_emails)
    print(emails)
    # create a beutiful soup for the html document
    soup = BeautifulSoup(response.text, 'lxml')

使用 Python 生成随机电子邮件地址:

from faker import Faker

faker = Faker()

for i in range(12):
    print(f'{faker.email()}')

【讨论】:

  • 非常感谢。那好极了。但是当尝试使用此站点https://www.randomlists.com/email-addresses 时,它不起作用。
  • 这是因为该网站在网站加载后使用 javascript 加载电子邮件地址。所以在爬取网页的时候,地址还没有出现在页面中,爬虫是找不到的。您可以使用我的其他答案使用 Python 生成随机电子邮件地址。
  • 非常感谢。有没有运气来处理 JavaScript ..?因为目标是从任何有或没有 JavaScript 的页面获取电子邮件。
猜你喜欢
  • 2019-07-04
  • 1970-01-01
  • 2013-01-24
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-08-11
  • 2011-11-30
相关资源
最近更新 更多