【问题标题】:Extracting sibling text nodes using Beautiful Soup使用 Beautiful Soup 提取兄弟文本节点
【发布时间】:2017-02-20 01:30:37
【问题描述】:

我正在尝试使用漂亮的汤来获取某些文本,但我不知道如何获取 /strong 标记后的文本。我找到了我正在寻找的内容,但只想要某些元素。

res = requests.get('http://www.fangraphs.com/statss.aspx?playerid=10155&position=OF')
res.raise_for_status()
soup = bs4.BeautifulSoup(res.text, "lxml")
gamescore = soup.select('#content > table > tr > td > table > tr > td > div')

输出: 出生日期: 1991 年 8 月 7 日(25 岁、6 米、12 天)击球次数/投掷次数: R/R

是否可以仅从中获取生日和 R/R?

【问题讨论】:

    标签: python beautifulsoup


    【解决方案1】:

    您可以根据文本选择<strong> 元素,然后使用next_sibling property 选择相邻的兄弟节点。

    birthday = soup.find('strong', text='Birthdate:').next_sibling.strip()
    gamescore = soup.find('strong', text='Bats/Throws:').next_sibling.strip()
    

    输出:

    > print(birthday, gamescore)
    > 8/7/1991 (25 y, 6 m, 12 d) R/R
    

    如果您想选择每个 <strong> 元素及其下一个兄弟节点,那么您可以使用以下内容:

    elements = soup.select('#content > table table div > strong')
    
    for element in elements:
        print(element.text, element.next_sibling)
    

    输出:

    > Birthdate:  8/7/1991 (25 y, 6 m, 12 d)     
    > Bats/Throws:  R/R     
    > Height/Weight:  6-1/235     
    > Position:  OF
    > Contract:
    

    【讨论】:

    • 那是完美的。谢谢
    猜你喜欢
    • 2014-04-23
    • 2018-08-17
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-11-01
    • 2011-11-03
    • 2018-09-06
    相关资源
    最近更新 更多