【发布时间】:2011-06-16 06:43:57
【问题描述】:
我正在尝试抓取投标网站的内容,但无法获取网站的完整页面。我在 xulrunner 上使用 crowbar 首先获取页面(因为 ajax 以惰性方式加载某些元素),然后从文件中抓取。 但是在 birivals 网站的主页上,即使本地文件格式正确,这也会失败。 jSoup 似乎只是在 html 代码中间以“...”字符结尾。 如果有人以前遇到过这种情况,请帮助。 为 [this link] 调用以下代码。
File f = new File(projectLocation+logFile+"bidrivalsHome");
try {
f.createNewFile();
log.warn("Trying to fetch mainpage through a console.");
WinRedirect.redirect(projectLocation+"Curl.exe -s --data \"url="+website+"&delay="+timeDelay+"\" http://127.0.0.1:10000", projectLocation, logFile+"bidrivalsHome");
} catch (Exception e) {
e.printStackTrace();
log.warn("Error in fetching the nameList", e);
}
Document doc = new Document("");
try {
doc = Jsoup.parse(f, "UTF-8", website);
} catch (IOException e1) {
System.out.println("Error while parsing the document.");
e1.printStackTrace();
log.warn("Error in parsing homepage", e1);
}
【问题讨论】:
-
你能发布你正在使用的生成
...的代码吗? -
添加了代码。此外,同样的事情通过 jSoup.connect(url).get()
-
@submit:但是在这里您已经构建了文档。 ... 究竟出现在哪里?
-
在文档对象内部。我希望它具有完整的 html 内容,但奇怪的是在页面中途失败并以 ellipse 结尾。因此,无法访问各种 div 和其他元素。
-
您正在抓取的页面被停放了?此外,如果这是您要抓取的网站,其中大部分是由 JavaScript 加载的,jsoup 无法处理。
标签: java web-scraping jsoup