【问题标题】:Extract emails from a data dump using Python使用 Python 从数据转储中提取电子邮件
【发布时间】:2016-06-07 15:56:23
【问题描述】:

我有一个数据转储,我试图从中提取所有电子邮件。

这是我使用 BeautifulSoup 编写的代码

import urllib2
import re
from bs4 import BeautifulSoup
url = urllib2.urlopen("file:///users/home/Desktop/emails.html").read()
soup = BeautifulSoup(url)
email = raw_input(soup)
match = re.findall(r'<(.*?)>', email)
if match:
    print match

示例数据转储

<tr><td><a href="http://abc.gov.com/comments/24-April/file.html">for educational purposes only</a></td>
<td>7418681641 &lt;sampleemail@gmail.com&gt;</td>
<td>advqos@abc.gov.com</td>
<td nowrap="">24-04-2015 10.31</td>
<td align="center">&nbsp;</td></tr>
<tr><td><a href="http://abc.gov.com/comments/24-April/test.html">no_subject</a></td>
<td>John &lt;someemail@gmail.com&gt;</td>
<td>advqos@abc.gov.com</td>
<td nowrap="">24-04-2015 11.28</td>
<td align="center">&nbsp;</td></tr>
<tr><td><a href="http://abc.gov.com/comments/24-April/test.html">something</a></td>
<td>Mark &lt;123random@gmail.com&gt;</td>
<td>test@abc.gov.com</td>
<td nowrap="">24-04-2015 11.28</td>
<td align="center">&nbsp;</td></tr>
<tr><td><a href="http://abc.gov.com/comments/24-April/abc.html">some data</a></td>

我可以清楚地看到电子邮件列在&amp;lt; 和&amp;gt; 标记之间。我正在尝试使用正则表达式来识别所有电子邮件并打印它们。但是,在执行时,不是只提取电子邮件(每行一封电子邮件),而是打印整个文件。

我该如何解决这个问题?

【问题讨论】:

  • 我完全看不懂你的代码。为什么要使用urllib2 打开本地文件?只需使用with open("/path/to/file.html") as f: soup = BeautifulSoup(f)。接下来,您希望raw_input(soup) 做什么?最后,为什么刚开始使用 HTML 解析器时要对文本进行正则表达式搜索?
  • @MattDMo:啊,是的,先生。可以简单地打开它。不知道 raw_input 接受用户的输入。我假设它将把汤变量解析成一个字符串。如果没有 raw_input 行,我收到一条错误消息,指出 re.findall 函数需要一个字符串作为字符串中的第二个参数

标签: python parsing web-scraping beautifulsoup


【解决方案1】:

你的例子确实有效

re.findall(r'\&lt;(.*?)\&gt;',your_data_bump)=
['sampleemail@gmail.com', 'someemail@gmail.com', '123random@gmail.com']

【讨论】:

  • 感谢这确实有效。我只是将这一行更改为 match = re.findall(r'
【解决方案2】:

假设您的数据转储位于名为 text.txt 的文本文件中:

import re
# Make sure the text file is in the same folder as the python file.
with open('text.txt','r') as f:
    matches = re.findall(r'&lt;(.+?)&gt;',f.read())
print('\n'.join(matches))

【讨论】:

    【解决方案3】:

    您可以使用BeautifulSoup 的find_all 方法解析到您要查找的标签。这是代码。 (我已将示例文件存储为a.html)

    from bs4 import BeautifulSoup
    url = open("a.html",'r').read()
    soup = BeautifulSoup(url)
    rows = soup.find_all('tr') # find all rows using tag 'tr'
    for row in rows:
        cols = row.find_all('td')  # find all columns using 'td' tag
        if len(cols)>1:
            email_id_string = cols[1].text # get the text of second element of list (contains email id element)
            email_id = email_id_string[ email_id_string.find("<")+1 : email_id_string.find(">") ] (get only the email id between < and > )
            print email_id
    

    【讨论】:

    • 这行不通,因为有很多 &lt;td&gt; 元素不包含电子邮件地址。
    • 如果电子邮件 ID 存在,则它作为第二列存在,因此我使用 if condition 进行了检查
    • 不,你没有。您只需检查每个&lt;tr&gt; 是否有多个&lt;td&gt; 元素,如果是,则取第二个 td 并假设它包含一封电子邮件,对于任意 HTML,这不是一个有效的假设。 OP 发布了一个非常简单的示例,而我认为实际数据的结构不如这个。您的解决方案需要比目前更强大。
    • 是的,你是对的。该解决方案完全基于 OP 发布的内容。我已经使用给定的示例 HTML 文件引用了这个解决方案。但是,OP 仍然可以从中找到一种方法,如何使用 BeautifulSoup 的简单迭代进一步进行而不涉及更多复杂性
    猜你喜欢
    • 1970-01-01
    • 2015-07-09
    • 2016-01-11
    • 2018-11-10
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-01-24
    • 1970-01-01
    相关资源
    最近更新 更多