【问题标题】:Extract links if anchor text contains keyword [duplicate]如果锚文本包含关键字,则提取链接[重复]
【发布时间】:2014-07-28 05:34:07
【问题描述】:
import BeautifulSoup

html = """
<html><head></head>
<body>
<a href='http://www.gurletins.com'>My HomePage</a>
<a href='http://www.gurletins.com/sections'>Sections</a>
</body>
</html>
"""

soup = BeautifulSoup.BeautifulSoup(html)

现在我想获取包含关键字Home的链接

谁能告诉我如何使用 BeautifulSoup 做到这一点?

【问题讨论】:

    标签: python web-scraping beautifulsoup


    【解决方案1】:
    html = """
    <html><head></head>
    <body>
    <a href='http://www.gurletins.com'>My HomePage</a>
    <a href='http://www.gurletins.com/sections'>Sections</a>
    </body>
    </html>
    """
    from bs4 import BeautifulSoup
    soup = BeautifulSoup(html)
    
    for i in soup.find_all("a"):
        if "HOME" in str(i).split(">")[1].upper():
            print i["href"]
    http://www.gurletins.com
    

    【讨论】:

      【解决方案2】:

      有更好的方法。在text 参数中传递正则表达式:

      import re
      from bs4 import BeautifulSoup
      
      html = """
      <html><head></head>
      <body>
      <a href='http://www.gurletins.com'>My HomePage</a>
      <a href='http://www.gurletins.com/sections'>Sections</a>
      </body>
      </html>
      """
      
      soup = BeautifulSoup(html)
      for a in soup.find_all("a", text=re.compile('Home')):
          print a['href']
      

      打印:

      http://www.gurletins.com
      

      请注意,默认情况下它区分大小写。如果您需要使其不敏感,请将re.IGNORECASE 标志传递给re.compile():

      re.compile('Home', re.IGNORECASE)
      

      演示:

      >>> import re
      >>> from bs4 import BeautifulSoup
      >>> 
      >>> html = """
      ... <html><head></head>
      ... <body>
      ... <a href='http://www.gurletins.com'>My HomePage</a>
      ... <a href='http://www.gurletins.com/sections'>Sections</a>
      ... <a href='http://www.gurletins.com/home'>So nice to be home</a>
      ... </body>
      ... </html>
      ... """
      >>> 
      >>> soup = BeautifulSoup(html)
      >>> for a in soup.find_all("a", text=re.compile('Home', re.IGNORECASE)):
      ...     print a['href']
      ... 
      http://www.gurletins.com
      http://www.gurletins.com/home
      

      【讨论】:

      • 不区分大小写吗?
      • 为什么需要重新搜索单个单词?
      • @PadraicCunningham 因为它可能在文本中的任何位置。省略 re.compile() 将有助于仅找到完全匹配。
      • 链接中的 home 也会匹配吗?
      • @PadraicCunningham 单个小写?默认情况下,它是敏感的,但由re.IGNORECASE 标志控制。
      猜你喜欢
      • 1970-01-01
      • 2013-01-29
      • 2017-08-03
      • 1970-01-01
      • 1970-01-01
      • 2016-09-29
      • 1970-01-01
      • 2021-08-13
      • 1970-01-01
      相关资源
      最近更新 更多