【问题标题】:How Do I get the Text Only From This Class ID With Span in Middle?如何仅从此类 ID 中获取文本,其中 Span 位于中间?
【发布时间】:2021-10-23 20:42:03
【问题描述】:

下面的文本返回该类的 div,但我希望将季度和时间分开。

我试图做 .text,但它给了我属性错误。由于季度文本的一部分由跨度标签分隔,我将如何获取文本?例如...第三季度 x:xx 看起来像:

div

“3”

“x:xx”

div

import pandas as pd
from bs4 import BeautifulSoup
import requests
import lxml


cbs_scores = requests.get('https://www.cbssports.com/college-football/scoreboard/', 'r').text
soup = BeautifulSoup(cbs_scores, 'lxml')

time_left = soup.find_all('div', class_ = 'game-status emphasis')

print(time_left)

【问题讨论】:

    标签: python-3.x web-scraping


    【解决方案1】:

    我得到以下输出:

    import pandas as pd
    from bs4 import BeautifulSoup
    import requests
    import lxml
    
    
    cbs_scores = requests.get('https://www.cbssports.com/college-football/scoreboard/', 'r').text
    soup = BeautifulSoup(cbs_scores, 'lxml')
    
    for  time_left in soup.find_all('div', class_ = 'game-status emphasis'):
        print(time_left.get_text())
    
                
    

    输出:

                4th 8:49
    
    
                3rd 7:52
    
    
                3rd 2:19
    
    
                3rd 7:55
    
    
                3rd 3:09
    
    
                3rd 6:37
    
    
                3rd 5:31
    
    
                End 3rd 
    
    
                4th 12:15
    
    
                3rd 4:30
    
    
                3rd 11:27
    
    
                3rd 7:40
    
    
                3rd 11:54
    
    
                3rd 12:45
    
    
                End 2nd
    
    
                End 2nd
    

    【讨论】:

    • 欣赏它。我尝试了一个 for 循环,但使用了 .text 而不是 get_text()。这很好用,再次感谢!
    猜你喜欢
    • 1970-01-01
    • 2021-07-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-07-04
    • 1970-01-01
    • 1970-01-01
    • 2021-06-13
    相关资源
    最近更新 更多