【发布时间】:2016-03-23 20:14:04
【问题描述】:
考虑以下网址http://www.google.com/url?rct=j&sa=t&url=http://www.ksat.com/news/father-of-woman-killed-in-memorial-day-floods-testifies-for-better-flood-warnings&ct=ga&cd=CAIyHWU3NmVhMGQ0NWQ3MmRmY2I6Y29tOmVuOlVTOlJM&usg=AFQjCNE_8XwECqkmyPIMzcSxCDh2hP16wQ。当我将此url 传递给JSOUP 时,html 内容不准确。但是当我在浏览器中打开这个网址时,它会重定向到http://www.ksat.com/news/father-of-woman-killed-in-memorial-day-floods-testifies-for-better-flood-warnings。
然后,我将这个url 传递给jsoup,现在我得到了准确的html 内容。
如何从第一个 url 中获取准确的 html 内容??
我尝试了很多选择
Response response = Jsoup.connect(url).followRedirects(true).timeout(timeOut*1000).userAgent(userAgent).execute();
int status = response.statusCode();
if (status == HttpURLConnection.HTTP_MOVED_TEMP || status == HttpURLConnection.HTTP_MOVED_PERM || status == HttpURLConnection.HTTP_SEE_OTHER) {
redirectUrl = response.header("location");
response = Jsoup.connect(redirectUrl).followRedirects(false).timeout(timeOut*1000).userAgent(userAgent).execute();
}
Document doc=response.parse();
我尝试了很多user agents、.referrer("http://google.com") 选项等。
我目前使用的是jsoup 1.8.3 版。
【问题讨论】:
标签: java url html-parsing jsoup