【发布时间】:2017-12-14 02:11:18
【问题描述】:
我在尝试将正则表达式与漂亮的汤一起使用时遇到了一点麻烦。
我的html如下:
[<strong>See the full calendar</strong>, <strong>See all events</strong>, <strong>See all committee meetings</strong>, <strong>526 spaces</strong>, <strong>89 spaces</strong>, <strong>53 spaces</strong>, <strong>154 spaces</strong>, <strong>194 spaces</strong>, <strong>See all news releases</strong>]
[<strong>See the full calendar</strong>, <strong>See all events</strong>, <strong>See all committee meetings</strong>, <strong>526 spaces</strong>, <strong>89 spaces</strong>, <strong>53 spaces</strong>, <strong>154 spaces</strong>, <strong>194 spaces</strong>, <strong>See all news releases</strong>]
我想要的只是强标签之间的空格数。
我尝试过使用:
print(soup.find_all(re.compile("\d\d\d\s[a-zA-Z]{6}|(strong)")))
但是,这将返回 print(soup.find_all('strong')) 所做的一切。
有人知道我哪里出错了吗?
【问题讨论】:
-
谢谢!试过
print(len(soup.find_all('strong').split())-1)但得到AttributeError: 'ResultSet' object has no attribute 'split'- 有什么想法吗? @Ludisposed -
你想要所有空格的总和还是每个强标签都应该有它自己的空格计数器?
-
最终目标是将其导出为 csv,因此每个“x 个空格”需要是每一行的单独记录
标签: python regex web-scraping beautifulsoup