【发布时间】:2017-09-29 23:54:23
【问题描述】:
这是我的代码:
import requests
find_doctors_url = 'http://www.americhoice.com/find_doctor/Ver2/results_doc.jsp'
payload ={"specialty":"ORSU","docproducts":"HOBD,HLOP","planNameDropDoc":"HOBD,HLOP","plan":"uhcwa","zip":"98122","zipradius":"10","findButton":"FIND DOCTOR","specialtyName":"ORTHOPAEDIC SURGERY"}
response = requests.get(find_doctors_url,params=payload)
print(response.url)
print(response.content)
当我打印 response.content 时,我收到的只是:
<!-- NEAADR0179 -Anil Kumar Vutikuri *** End-->
<!--BEGIN SETTING HEADERS TO NO CACHE-->
<!--END SETTING HEADERS TO NO CACHE-->
<!--SET SESSION VALUES FROM URL PARAMETERS-->
<!--END SET SESSION VALUES FROM URL PARAMETERS-->
当您导航到: 查看源代码:http://www.americhoice.com/find_doctor/Ver2/results_doc.jsp
但是,当您导航到由 response.url 生成的 url 时,我正在寻求返回收到的完整 html
问题似乎是请求没有正确发送 GET 查询参数
我尝试过的事情(不成功): 1)请求完整的 URL(编码)而不是使用 params 字典 2)使用 urllib3 库而不是 Requests 库
【问题讨论】:
标签: python web-scraping get python-requests urllib3