【发布时间】:2017-07-20 07:37:32
【问题描述】:
我只想提取此webpage 中标题为“受访者在说什么……”下的项目符号。
我可以用这段代码实现它:
import requests
URL = 'https://www.instituteforsupplymanagement.org/about/MediaRoom/newsreleasedetail.cfm?ItemNumber=30655&SSO=1'
r = requests.get(URL)
page = r.text
from bs4 import BeautifulSoup
soup = BeautifulSoup(page, 'lxml')
strong_el = soup.find('strong',text='WHAT RESPONDENTS ARE SAYING …')
strong_el.find_all_next('li')[9]
但这里的问题是我必须知道列出了多少个要点(在这种情况下有 10 个。因此,它返回 [9] 之前的有效值)。即使不知道列出了多少要点,提取所有要点的最佳方法是什么?另外,我只需要文本而不需要 html。
【问题讨论】:
标签: python regex parsing beautifulsoup python-requests