【问题标题】:Python 3 - Extract content between <td></td> [duplicate]Python 3 - 提取 <td></td> 之间的内容 [重复]
【发布时间】:2018-01-27 20:55:24
【问题描述】:
from bs4 import BeautifulSoup
import re

data = open('C:\folder')
soup = BeautifulSoup(data, 'html.parser')
emails = soup.find_all('td', text = re.compile('@'))

for line in emails:
   print(line)

我有上面的脚本,它在 Python 2.7 中完美运行,并带有 Beautifulsoup,用于在 HTML 文件中的多个之间提取内容。然而,当我在 Python 3.6.4 中运行相同的脚本时,我得到以下结果:

<td>xxx@xxx.com</td>
<td>xxx@xxx.com</td>

我想要没有 TD 的内容...

为什么在 Python 3 中会发生这种情况?

【问题讨论】:

    标签: python python-3.x beautifulsoup


    【解决方案1】:

    我找到了答案……

    from bs4 import BeautifulSoup
    import re
    
    data = open('C:\folder')
    soup = BeautifulSoup(data, 'html.parser') #Lade till html.parser
    emails = soup.find_all('td', text = re.compile('@'))
    
    for td in emails:
       print(td.get_text())
    

    仔细看最后两行:)

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-09-17
      • 1970-01-01
      • 2013-01-08
      • 2016-09-10
      • 1970-01-01
      相关资源
      最近更新 更多