【问题标题】:node.js \ why do I get RangeError: Maximum call stack size exceedednode.js \ 为什么我得到 RangeError: Maximum call stack size exceeded
【发布时间】:2015-07-26 18:41:18
【问题描述】:

以下程序的目的是抓取 CNN,并将其所有文本写入单个文件(使用几个第三方)

我明白了

RangeError: Maximum call stack size exceeded

如何解决这个问题,我该如何绕过它?有没有办法可以“释放”内存?以及如何?

//----------Configuration--------------

var startingUrl = "http://cnn.com"; //keep the http\https or www prefix
var crawlingDepth = "50";
var outputFileName = "cnn.txt";

//-------------------------------------

var Crawler = require("js-crawler");
var sanitizeHtml = require('sanitize-html');
var htmlToText = require('html-to-text');
var fs = require('fs');

var index = 0;

new Crawler().configure({depth: crawlingDepth})
  .crawl(startingUrl, function onSuccess(page) {

  var text = htmlToText.fromString(page.body, {
        wordwrap: false,
        hideLinkHrefIfSameAsText: true,
        ignoreHref: true,
        ignoreImage: true
    });

    index++;
    console.log(index + " pages were crawled"); 
    fs.appendFile(outputFileName, text, function (err) {
        if (err) {
            console.log(err);
        };
        console.log('It\'s saved! in same location.');
    }); 
  });

【问题讨论】:

  • 您是否尝试降低crawlingDepth?我猜你只是爬到了很多都保存在内存中的页面。
  • 自然地它适用于较低的深度。所以你暗示 js-crawler 递归地工作?也许我会尝试其他爬虫..
  • 是的,它是递归的。检查Crawler.prototype._crawlUrl。

标签: node.js web-crawler out-of-memory html-to-text


【解决方案1】:

1) 这是递归深度的问题。

2) 必须避免:

  • 在每个深度级别由环路电流级别中的链接遍历(第一级别是一个主要参考);

  • 使用当前页面的“Crawler.prototype._getAllUrls”链接进行访问,如果这些链接尚未处理 - 循环访问它们;

3) 唯一的概念:

var Urls = [ ["http://cnn.com/"] ]; // What we crawling
var crawledUrls = {}; // Check if already crawled
var crawlingDepth = 3;
var depth = 0; // Current depth
var index = 0; // Current index
var Crawler = require("js-crawler");

function crawling() {
  console.log(depth, index, Urls[depth][index]);

  // Prepare next level
  if (typeof Urls[depth+1] === "undefined") Urls.push([]);

  // Already crawled flag
  crawledUrls[ Urls[depth][index] ] = true;

    new Crawler().configure({depth: 1}).crawl({
        url: Urls[depth][index],
        success: function(page) {
            // Do some with crawled page

            // Collect urls at crawled page
            var urls = Crawler.prototype._getAllUrls( page.url, page.body );
            for(var j=0; j<urls.length; j++) {
                // Check same domain and now crawled yet
                if ( typeof crawledUrls[urls[j]] === "undefined"
                     && urls[j].indexOf(Urls[0][0])===0 ) {
                    Urls[depth+1].push(urls[j]);
                }
            }
        },
        failure: function(page) {
        },
        finished: function(crawled) {
          index++;
          if (index<Urls[depth].length) {
            setTimeout(crawling,0);
          } else {
            depth++;
            index = 0;
            if (depth<crawlingDepth) {
              setTimeout(crawling,0);
            } else {
              // Finished
            }
        }
        }
    });
}

crawling();

【讨论】:

  • 谢谢 - 不确定我是否在关注 - 你是否建议我编写一些代码来克服这个问题?喜欢用一些逻辑处理我自己的 allUrls 变量吗?
  • @user1025852 更新:)
  • 哦,这太棒了 :) - 出于某种原因,大约 2K 页,上面的代码只是挂起......试图看看是什么持有它......
  • 只是摆脱递归的概念 :) 其实这个库客观上不适合这么多的页面。如果您想要一个不太严重的决定,则有必要根据数据库来查看决定。也许像弹性搜索。
猜你喜欢
  • 2015-10-19
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-12-25
  • 1970-01-01
相关资源
最近更新 更多