【发布时间】:2022-02-01 03:24:56
【问题描述】:
我正在从网页中提取表格和表格标题元素。表格元素已被提取,没有任何问题。但是,我无法将 h2 类提取到单独的字符串中。我可以将所有内容都导入为 beautifulsoup 对象,也可以导入为一个包含所有 h2 元素的长字符串。如何将元素作为单独的字符串对象提取到表格或列表中?
scr = 'https://tv.varsity.com/results/7361971-2022-spirit-unlimited-battle-at-the-
boardwalk-atlantic-city-grand-ntls/31220'
scr1 = requests.get(scr)
soup = BeautifulSoup(scr1.text, "html.parser")
sp3 = soup.find(class_="full-content").find_all("h2")
这是我目前尝试过的两种方法。
comp = pd.DataFrame(sp3[0], dtype=str)
div1a = div.drop(div.iloc[0].name)
div2a = div1a.drop(div1a.iloc[0].name)
也使用 for 循环
data = []
for a in soup.find(class_="full-content").find_all("h2"):
a = str(a.text)
data.append(a)
x = ",".join(map(str, data))
print(x)
感谢您的帮助!
【问题讨论】:
标签: python beautifulsoup css-selectors