【问题标题】:Accessing a website in python在 python 中访问网站
【发布时间】:2015-10-22 22:33:03
【问题描述】:

我正在尝试使用 python 获取网站上的所有网址。目前我只是将网站 html 复制到 python 程序中,然后使用代码提取所有 url。有没有一种方法可以直接从网络上执行此操作,而无需复制整个 html?

【问题讨论】:

  • 你用的是什么库?硒?碎片? urllib?
  • 我不确定这些意味着我对编程很陌生

标签: python url web


【解决方案1】:

在 Python 2 中,您可以使用 urllib2.urlopen:

import urllib2
response = urllib2.urlopen('http://python.org/')
html = response.read()

在 Python 3 中,您可以使用 urllib.request.urlopen:

import urllib.request
with urllib.request.urlopen('http://python.org/') as response:
    html = response.read()

如果您必须执行更复杂的任务,例如身份验证或传递参数,我建议您查看requests 库。

【讨论】:

    【解决方案2】:

    最直接的可能是urllib.urlopen,如果你使用的是python2,或者urllib.request.urlopen,如果你使用的是python3(你必须首先做import urllib或import urllib.request)。这样您就可以得到一个类似对象的文件,您可以从中读取(即f.read())html 文档。

    python 2 示例:

    import urllib
    
    f = urlopen("http://stackoverflow.com")
    
    http_document = f.read()
    f.close()
    

    好消息是您似乎已经完成了分析 html 文档中链接的困难部分。

    【讨论】:

      【解决方案3】:

      您可能想要使用 bs4(BeautifulSoup) 库。

      Beautiful Soup 是一个 Python 库,用于从 HTML 和 XML 文件中提取数据。

      您可以在 cmd 行使用 followig 命令下载 bs4。 pip install BeautifulSoup4

      import urllib2
      import urlparse
      from bs4 import BeautifulSoup
      
      url = "http://www.google.com"
      response = urllib2.urlopen(url)
      content = response.read()
      
      soup = BeautifulSoup(content, "html.parser")
      for link in soup.find_all('a', href=True):
          print urlparse.urljoin(url, link['href'])
      

      【讨论】:

        【解决方案4】:

        您可以简单地使用requests 和BeautifulSoup 的组合。

        • 首先使用requests 发出HTTP 请求以获取HTML 内容。你会得到一个 Python 字符串,你可以随意操作它。
        • 获取 HTML 内容字符串并将其提供给 BeautifulSoup,它已完成提取 DOM 的所有工作,并获取所有 URL,即 <a> 元素。

        这是一个如何从 StackOverflow 获取所有链接的示例:

        import requests
        from bs4 import BeautifulSoup, SoupStrainer
        
        response = requests.get('http://stackoverflow.com')
        html_str = response.text
        
        bs = BeautifulSoup(html_str, parseOnlyThese=SoupStrainer('a'))
        
        for a_element in bs:
            if a_element.has_attr('href'):
                print(a_element['href'])
        

        样本输出:

        /questions/tagged/facebook-javascript-sdk
        /questions/31743507/facebook-app-request-dialog-keep-loading-on-mobile-after-fb-login-called
        /users/3545752/user3545752
        /questions/31743506/get-nuspec-file-for-existing-nuget-package
        /questions/tagged/nuget
        ...
        

        【讨论】:

          猜你喜欢
          • 2018-02-15
          • 1970-01-01
          • 2014-01-04
          • 2022-12-14
          • 2013-05-24
          • 2014-11-13
          • 2018-03-10
          • 2022-11-01
          • 1970-01-01
          相关资源
          最近更新 更多