【问题标题】:extract e-mails from multiple pages in a website and list it从网站的多个页面中提取电子邮件并列出
【发布时间】:2019-08-11 02:07:10
【问题描述】:

我想使用 python 从一个展览网站中提取参展商的电子邮件。该页面包含参展商的超文本。点击参展商名称后,您将找到包含其电子邮件的参展商资料。

您可以在这里找到该网站:

https://www.medica-tradefair.com/cgi-bin/md_medica/lib/pub/tt.cgi/Exhibitor_index_A-Z.html?oid=80398&lang=2&ticket=g_u_e_s_t

请问我该如何使用 python 做到这一点? 提前谢谢你

【问题讨论】:

  • 请向我们展示您的代码,以便我们提供帮助。
  • 有很多项目可以帮助您抓取页面。您可以为此使用硒。
  • 在此上下文中不包含代码的问题应该因为过于宽泛而被关闭。请添加您当前的编码尝试和研究。

标签: python web-scraping scrapy python-requests web-crawler


【解决方案1】:

您可以获取所有参展商的链接,然后遍历这些链接并提取每个参展商的电子邮件:

import requests
import bs4


url = 'https://www.medica-tradefair.com/cgi-bin/md_medica/lib/pub/tt.cgi/Exhibitor_index_A-Z.html?oid=80398&lang=2&ticket=g_u_e_s_t'

response = requests.get(url)

soup = bs4.BeautifulSoup(response.text, 'html.parser')

links = soup.find_all('a', href=True)
exhibitor_links = ['https://www.medica-tradefair.com'+link['href'] for link in links if 'vis/v1/en/exhibitors' in link['href'] ]
exhibitor_links = list(set(exhibitor_links))

for link in exhibitor_links:
    response = requests.get(link)
    soup = bs4.BeautifulSoup(response.text, 'html.parser')

    name = soup.find('h1',{'itemprop':'name'}).text
    try:
        email = soup.find('a', {'itemprop':'email'}).text
    except:
        email = 'N/A'

    print('Name: %s\tEmail: %s' %(name, email))

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-11-30
    • 1970-01-01
    • 2017-09-16
    • 1970-01-01
    • 2021-12-23
    • 2022-11-15
    相关资源
    最近更新 更多