【问题标题】:Trying to get the texts between this tag but getting an empty list试图获取此标签之间的文本但得到一个空列表
【发布时间】:2019-07-18 10:27:19
【问题描述】:

\试图从此 html 中获取文本 A Plus 和 Computers:

<div class="u-space-t1">
        <h1 class="biz-page-title embossed-text-white shortenough">A Plus</h1>
        <div class="u-inline-block">
            <h1 class="biz-page-title\ embossed-text-white\ shortenough">Computers</h1>
            <div class="u-inline-block"> 

所以我试着得到这样的文字:

c = soup.findAll('h1',{"class":"biz-page-title embossed-text-white shortenough"})

print(c)

但是我得到一个空列表

我也试过这样做:

c = soup.find('div', class_='u-inline-block').h1

我得到一个未找到的“Nonetype”对象。

【问题讨论】:

    标签: html python-3.x python-2.7 web-scraping beautifulsoup


    【解决方案1】:

    这样做。

    texts = soup.select("div > h1, div > div > h1")
    for text in texts:
        print(text.text)
    

    “A Plus”和“计算机”将会问世。

    【讨论】:

    • 感谢您的回复。但是,我仍在打印其他内容。有没有办法更具体地查找文本。例如使用 class 属性?
    【解决方案2】:

    试试这个:

    html = """
    <div class="u-space-t1">
            <h1 class="biz-page-title embossed-text-white shortenough">A Plus</h1>
            <div class="u-inline-block">
                <h1 class="biz-page-title\ embossed-text-white\ shortenough">Computers</h1>
                <div class="u-inline-block"> 
    """
    
    soup = bs4(html, 'lxml')
    for i in soup.find_all('h1'):
        print(i.text)
    

    输出:

    A Plus
    Computers
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-06-26
      • 1970-01-01
      相关资源
      最近更新 更多