【问题标题】:How to get exact page content in wget if error code is 404如果错误代码为 404,如何在 wget 中获取准确的页面内容
【发布时间】:2015-05-27 06:57:26
【问题描述】:

我有两个 url 一个是工作 url 另一个是页面已删除 url。工作 url 很好,但对于页面删除 url 而不是获取确切的页面内容 wget 接收 404

工作网址

import os
def curl(url):
    data = os.popen('wget -qO- %s '% url).read()
    print (url)
    print (len(data))
    #print (data)

curl("https://www.reverbnation.com/artist_41/bio")

输出:

https://www.reverbnation.com/artist_41/bio
80067

页面删除网址

import os
def curl(url):
    data = os.popen('wget -qO- %s '% url).read()
    print (url)
    print (len(data))
    #print (data)

curl("https://www.reverbnation.com/artist_42/bio")

输出:

https://www.reverbnation.com/artist_42/bio
0

我的长度为 0,但实时页面中有一些内容

如何在 wget 或 curl 中接收准确的内容

【问题讨论】:

    标签: python-3.x curl web-scraping wget


    【解决方案1】:

    wget 有一个名为“--content-on-error”的开关:

    --content-on-error
               If this is set to on, wget will not skip the content when the server responds with a http status code that indicates error.
    

    所以只需将其添加到您的代码中,您也将拥有 404 页面的“内容”:

    import os
    def curl(url):
        data = os.popen('wget --content-on-error -qO- %s '% url).read()
        print (url)
        print (len(data))
        #print (data)
    
    curl("https://www.reverbnation.com/artist_42/bio")
    

    【讨论】:

      猜你喜欢
      • 2019-04-07
      • 2012-10-12
      • 2011-11-24
      • 1970-01-01
      • 2021-06-09
      • 1970-01-01
      • 2021-09-07
      • 2013-08-11
      • 1970-01-01
      相关资源
      最近更新 更多