【问题标题】:Using RegEx to identify emails from Beautiful Soup使用 RegEx 识别来自 Beautiful Soup 的电子邮件
【发布时间】:2020-04-04 02:18:15
【问题描述】:

我是一名初学者,正在开发一个可以从给定网站上抓取电子邮件的程序。代码如下:

import requests, bs4, re
print('Fetching Website...')
res = requests.get('https://examplewebsite.com')
res.raise_for_status()
soup = bs4.BeautifulSoup(res.text, 'html.parser')
type(soup)

my_list = []
for link in soup.find_all('a'):
    my_list.append(link.get('href'))

emailregex = re.compile(r'''(
    [a-zA-Z0-9._%+-:]+
    @
    [a-zA-Z0-9.-]+
    \.[a-zA-Z]{2,4}
    )''', re.VERBOSE)

newlist = list(filter(emailregex.search, my_list))
print(newlist)

print('---Done---')

但是,当我运行代码时,我收到一个错误:“TypeError:预期的字符串或类似字节的对象”。我发现如果我这样做:

newlist = list(filter(emailregex.search, str(my_list)))
print(newlist)

错误会消失,但我的“新列表”不包含任何结果。我已验证“my_list”确实返回了预期结果列表。我发现如果我打印“my_list”并将其内容粘贴到一个新文件中,然后将其添加到列表中运行相同的代码,它就可以正常工作,所以我不相信它是正则表达式的问题。我认为这可能与“my_list”中的数据类型有关?我真的没有什么好主意,所以任何帮助都将不胜感激。

谢谢

【问题讨论】:

  • 你真的是来自anchor标签的extractingemails吗?
  • 您的电子邮件正则表达式真的很差,看看这些网站:TLD list; valid/invalid addresses; regex for RFC822 email address
  • @αԋɱҽԃ αмєяιcαη 是的,代码提取格式为“mailto:jdoe@gmail.com”的电子邮件(至少对于我一直在查看的特定网站)。我才学了几个星期,所以我确信有更好的方法来实现这一点。
  • @toto 谢谢你的帮助。我刚开始学习,所以我相信还有很大的改进空间。我会检查你的链接。

标签: python regex list beautifulsoup


【解决方案1】:

"TypeError: expected string or bytes-like object" 是因为my_list 不只包含字符串,但是str(my_list) 会将变量转换为大字符串

print(str(my_list)) # this is a string
print(type(str(my_list)))  # output: str

你需要把my_list的每一项都改成字符串,然后再试一次

my_list = list(map(str, my_list))
newlist = list(filter(emailregex.search, my_list))

【讨论】:

    【解决方案2】:
    import requests
    from bs4 import BeautifulSoup
    import re
    
    
    def main(url):
        r = requests.get(url)
        soup = BeautifulSoup(r.content, 'html.parser')
        target = "".join([item.get("href")
                          for item in soup.findAll("a", href=True)])
        matches = re.findall(
            r'''[a-zA-Z0-9._%+-:]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,4}''', re.VERBOSE, target)
        for match in matches:
            print(match)
    
    
    main("https://www.example.com")
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-04-14
      • 2023-04-10
      相关资源
      最近更新 更多