【问题标题】:Find the final redirected url using Python使用 Python 查找最终重定向的 url
【发布时间】:2019-12-16 16:04:02
【问题描述】:

我正在尝试使用 python 来查找 url 的最终重定向 URL。我尝试了stackoverflow答案中的各种解决方案,但对我没有任何帮助。我只得到原始网址。

具体来说,我尝试了requests、urllib2 和urlparse 库,但它们都没有按应有的方式工作。以下是我尝试的一些代码:

解决方案 1:

s = requests.session()
r = s.post('https://www.boots.com/search/10055096', allow_redirects=True)
print(r.history)
print(r.history[1].url)

结果:

[<Response [301]>, <Response [302]>]
https://www.boots.com/search/10055096

解决方案 2:

import urlparse
url = 'https://www.boots.com/search/10055096'
try:
    out = urlparse.parse_qs(urlparse.urlparse(url).query)['out'][0]
    print(out)
except Exception as e:
    print('not found')

结果: not found

解决方案 3:

import urllib2
def get_redirected_url(url):
    opener = urllib2.build_opener(urllib2.HTTPRedirectHandler)
    request = opener.open(url)
    return request.url
print(get_redirected_url('https://www.boots.com/search/10055096'))

结果:

HTTPError: HTTP Error 302: The HTTP server returned a redirect error that would lead to an infinite loop.
The last 30x error message was:
Found

下面的预期 URL 是最终的重定向页面,这就是我想要返回的内容。

原网址:https://www.boots.com/search/10055096

预期网址: https://www.boots.com/gillette-fusion5-razor-blades-4pk-10055096

解决方案 #1 是最接近的解决方案。至少它返回了 2 个响应,但第二个响应不是最终页面,看起来它是正在查看它的内容的加载页面。

【问题讨论】:

    标签: http url redirect web-scraping python-requests


    【解决方案1】:

    第一个请求返回一个包含用于更新站点的 JS 的 html 文件,并且 requests 不处理 Java 脚本。您可以使用

    找到更新后的链接
    import requests
    from bs4 import BeautifulSoup
    import re
    
    r = requests.get('https://www.boots.com/search/10055096')
    soup = BeautifulSoup(r.content,'html.parser')
    reg = soup.find('input',id='searchBoxText').findNext('script').contents[0]
    print(re.search(r'ht[\w\://\.-]+', reg).group())
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2014-02-22
      • 2011-01-23
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-12-07
      相关资源
      最近更新 更多