【问题标题】:BeautifulSoup object contents to stringBeautifulSoup 对象内容到字符串
【发布时间】:2022-02-01 03:24:56
【问题描述】:

我正在从网页中提取表格和表格标题元素。表格元素已被提取,没有任何问题。但是,我无法将 h2 类提取到单独的字符串中。我可以将所有内容都导入为 beautifulsoup 对象,也可以导入为一个包含所有 h2 元素的长字符串。如何将元素作为单独的字符串对象提取到表格或列表中?

scr = 'https://tv.varsity.com/results/7361971-2022-spirit-unlimited-battle-at-the- 
boardwalk-atlantic-city-grand-ntls/31220'
    
scr1 = requests.get(scr)
soup = BeautifulSoup(scr1.text, "html.parser")
sp3 = soup.find(class_="full-content").find_all("h2")

这是我目前尝试过的两种方法。

comp = pd.DataFrame(sp3[0], dtype=str)
div1a = div.drop(div.iloc[0].name)
div2a = div1a.drop(div1a.iloc[0].name)

也使用 for 循环

data = []
for a in soup.find(class_="full-content").find_all("h2"):
    a = str(a.text)
    data.append(a)

x = ",".join(map(str, data))
print(x)

感谢您的帮助!

【问题讨论】:

    标签: python beautifulsoup css-selectors


    【解决方案1】:

    您可以使用列表推导式获取列表中每个 h2 元素的文本,或使用 for 循环遍历 h2 元素。

    import requests
    from bs4 import BeautifulSoup
    
    url = 'https://tv.varsity.com/results/7361971-2022-spirit-unlimited-battle-at-the-boardwalk-atlantic-city-grand-ntls/31220'
    response = requests.get(url)
    soup = BeautifulSoup(response.text, "html.parser")
    sp3 = soup.find(class_="full-content").find_all("h2")
    headers = [elt.text for elt in sp3]
    print(headers)
    

    输出:

    ['2022 Spirit Unlimited: Battle at the Boardwalk Atlantic City Grand Ntls Nationals Results',
    'Level 5 & 6 Results', 'L5 Junior', ...
    ]
    

    【讨论】:

    • 感谢您的帮助!这对我来说似乎很有效。
    猜你喜欢
    • 1970-01-01
    • 2015-09-22
    • 1970-01-01
    • 2016-10-26
    • 1970-01-01
    • 2015-08-29
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多