【问题标题】:How to get all results in a http request python如何在http请求python中获取所有结果
【发布时间】:2017-05-18 11:55:46
【问题描述】:

我正在尝试从https://www.ncl.com/ 获取所有结果。发现请求一定是GET,发到这个链接:https://www.ncl.com/search_vacations 到目前为止,我得到了前 12 个结果,解析它们没有问题。问题是我找不到“更改”结果页面的方法。我得到了 499 个中的 12 个,我需要把它们都拿下来。我试过这样做https://www.ncl.com/search_vacations?current_page=1 并每次都增加它,但我每次都得到相同的(第一个)结果。尝试再次将 json 正文添加到请求 json = {"current_page": '1'} 中,但没有成功。 到目前为止,这是我的代码:

    import math
import requests

session = requests.session()
proxies = {'https': 'https://97.77.104.22:3128'}
headers = {
    "authority": "www.ncl.com",
    "method": "GET",
    "path": "/search_vacations",
    "scheme": "https",
    "accept": "application/json, text/plain, */*",
    "connection": "keep-alive",
    "referer": "https://www.ncl.com",
    "cookie": "AkaUTrackingID=5D33489F106C004C18DFF0A6C79B44FD; AkaSTrackingID=F942E1903C8B5868628CF829225B6C0F; UrCapture=1d20f804-718a-e8ee-b1d8-d4f01150843f; BIGipServerpreprod2_www2.ncl.com_http=61515968.20480.0000; _gat_tealium_0=1; BIGipServerpreprod2_www.ncl.com_r4=1957341376.10275.0000; MP_COUNTRY=us; MP_LANG=en; mp__utma=35125182.281213660.1481488771.1481488771.1481488771.1; mp__utmc=35125182; mp__utmz=35125182.1481488771.1.1.utmccn=(direct)|utmcsr=(direct)|utmcmd=(none); utag_main=_st:1481490575797$ses_id:1481489633989%3Bexp-session; s_pers=%20s_fid%3D37513E254394AD66-1292924EC7FC34CB%7C1544560775848%3B%20s_nr%3D1481488775855-New%7C1484080775855%3B; s_sess=%20s_cc%3Dtrue%3B%20c%3DundefinedDirect%2520LoadDirect%2520Load%3B%20s_sq%3D%3B; _ga=GA1.2.969979116.1481488770; mp__utmb=35125182; NCL_LOCALE=en-US; SESS93afff5e686ba2a15ce72484c3a65b42=5ecffd6d110c231744267ee50e4eeb79; ak_location=US,NY,NEWYORK,501; Ncl_region=NY; optimizelyEndUserId=oeu1481488768465r0.23231006365903206",
    "Proxy-Authorization": "Basic QFRLLTVmZjIwN2YzLTlmOGUtNDk0MS05MjY2LTkxMjdiMTZlZTI5ZDpAVEstNWZmMjA3ZjMtOWY4ZS00OTQxLTkyNjYtOTEyN2IxNmVlMjlk"
}


def get_count():
    response = requests.get(
        "https://www.ncl.com/search_vacations?cruise=1&cruiseTour=0&cruiseHotel=0&cruiseHotelAir=0&flyCruise=0&numberOfGuests=4294953449&state=undefined&pageSize=10&currentPage=",
        proxies=proxies)
    tmpcruise_results = response.json()
    tmpline = tmpcruise_results['meta']
    total_record_count = tmpline['aggregate_record_count']
    return total_record_count


total_cruise_count = get_count()
total_page_count = math.ceil(int(total_cruise_count) / 10)
session.headers.update(headers)
cruises = []
page_counter = 1
while page_counter <= total_page_count:
    url = "https://www.ncl.com/search_vacations?current_page=" + str(page_counter) + ""
    page = requests.get(url, headers=headers, proxies=proxies)
    cruise_results = page.json()
    for line in cruise_results['results']:
        cruises.append(line)
        print(line)
    page_counter += 1
    print(cruise_results['pagination']["current_page"])
    print("----------")
print(len(cruises))

使用requests 和代理。任何想法如何做到这一点?

【问题讨论】:

  • 首先使用网络浏览器看看它在浏览器中是如何工作的。您可能需要 requests.Session() 才能轻松使用 cookie。
  • 尝试按照建议创建一个 Session(),然后记住将标头与您的请求一起发送(在标头中通常有 Cookie)
  • 我正在使用 firefox-developer 和 session。正如您在代码session.headers.update(headers) 中看到的那样。问题是当我收到回复时,回复中会显示current_page:1。这意味着我必须能够改变它们。到目前为止,即使在浏览器中我也找不到如何做到这一点。
  • 标题不包含任何关于页面或页码的内容。只有用户代理和位置信息。
  • 贴出了整个代码

标签: python json httprequest python-requests python-3.5


【解决方案1】:

该网站声称有 12264 个搜索结果(用于空白搜索),按 12 页组织。

搜索 url 带有一个参数 Nao,它似乎定义了搜索结果页的起始偏移量。

所以获取https://www.ncl.com/uk/en/search_vacations?Nao=45

应该得到一个包含 12 个搜索结果的“页面”,从结果 46 开始。

果然:

"pagination": {

    "starting_record": "46",
    "ending_record": "57",
    "current_page": "4",
    "start_page": "1",
    ...

因此,要对所有结果进行分页,请从 Nao = 0 开始,并为每次提取添加 12。

【讨论】:

  • 我去了website,在页面顶部的搜索框中输入了一个空白搜索。生成的search page 显示搜索结果的分页链接。检查分页链接显示Nao 在每个链接上都会增加,测试 API 是否响应相同的查询参数似乎是合理的。
猜你喜欢
  • 2018-12-15
  • 1970-01-01
  • 2019-04-03
  • 2016-10-14
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-09-15
  • 1970-01-01
相关资源
最近更新 更多