【问题标题】:Using promises to call a function after my first one is done?在我的第一个函数完成后使用 promises 调用函数?
【发布时间】:2016-12-08 18:28:56
【问题描述】:

我很难理解承诺。

我正在创建一个文件,该文件使用节点和 NPM 抓取网站,然后将数据记录到 CSV 文件中。现在我正在通过多次刮擦很好地收集数据,但我想在所有刮擦完成后调用写入 CSV 文件的函数。

有人可以告诉我如何创建一个承诺,该承诺将在调用 FileWrite 函数之前等到“scraper”函数中的所有刮擦完成?

现在我正在使用 request-promise 发出请求,然后对数据进行处理,但我对如何在多个请求发生后使 FileWrite 函数发生感到困惑。我尝试将 FileWrite 调用放在其中一个请求承诺中,但它们都在迭代多个要抓取的元素,我不希望文件多次写入。

'use strict';

//require NPM packages


//I chose to use request to make the http calls because it is very easy to use.
//This npm package also has recent updates, within the last 2 days.
//Lastly it has a huge number of downloads, this means it has a solid reputation in the community
var request = require('request');


//I chose to use cheerio to write the jquery for our node scraper,
//This package is very simple to use, and it was easy to write jQuery I was already familiar with,
//Cheerio also makes it simple for us to work with HTML elements on the server.
//Lastly, Cheerio is popular within the community, with continuous updates and a lot of downloads.
var cheerio = require('cheerio');

var rp = require('request-promise');

var fs = require('fs');


//I used the json2csv npm package because it was easy to implement into my code,
//This module also has frequent updates and heavy download activity.
//This is the most elegant package to download for simple translation of json objects to a CSV file format.
var json2csv = require('json2csv');




//Array for shirts JSON object for json2csv to write.
var ShirtProps = [];
var Counter = 0;
var homeURL = "http://www.shirts4mike.com/";


//start the scraper
scraper()


//Initial scrape of the home page, looking for shirts
function scraper () {

  //use the datafolderexists function to check if data is a directory
  if (!DataFolderExists('data')) {
    fs.mkdir('data');
  }
  //initial request of the home url to find links that may have shirts in them
rp(homeURL).then(function (html) {

    //use cheerio to load the HTML for scraping
    var $ = cheerio.load(html);
    //For every link with shirt in it iterate over the link and make a request.
    $("a[href*=shirt]").each(function() {


        //request promise 
        rp('http://www.shirts4mike.com/' + $(this).attr("href")).then(function (html) {
            Counter ++;
            //pass the html into the shirt data creator, so if it wound up scraping individual shirts from any of the links it adds it to the data object
            var $ = cheerio.load(html);
            //if the add to cart input exists, log the data to the shirtprops arary.
            if ($('input[value="Add to Cart"]').length) {
              var ShirtURL = $(this).find('a').attr('href');
              var time = new Date();
              //json array for json2csv
              var ShirtData = {
              Title: $('title').html(),
              Price: $('.price').html(),
              ImageURL: $('img').attr('src'),
              URL: homeURL + ShirtURL,
              Time: time.toString() 
              };
                ShirtProps.push(ShirtData);
                console.log(ShirtData);
            } else {
              //else we are on a products page, scrape those links for shirt data
                $('ul.products li').each(function() {
                var ShirtURL = $(this).find('a').attr('href');
                    rp('http://www.shirts4mike.com/' + ShirtURL).then(function (html){

                    var $ = cheerio.load(html);
                    var time = new Date();
                    var ShirtData = {
                    Title: $('title').html(),
                    Price: $('.price').html(),
                    ImageURL: $('img').attr('src'),
                    Url: homeURL + ShirtURL,
                    Time: time.toString()
                  };
                  ShirtProps.push(ShirtData);
                  console.log(ShirtData);

          }).catch(function(error) {
          console.error(error.message);
          console.error('Scrape failed from: ' + homeURL + 'blah2' + ' The site may be down, or your connection may need troubleshooting.');
          }); //end catch error
      }); //end products li each
              } //end else



    }).catch(function(error) {  //end rp
      console.error(error.message); //end if
  //tell the user in lamens terms why the scrape may have failed.
      console.error('Scrape failed from: ' + homeURL + 'blah' + ' The site may be down, or your connection may need troubleshooting.');
    }); //end catch error
  });  //end href each
    //one thing all shirts links have in common, they are contained in a div with class shirts, find the link to the shirts page based on this class.

    // //console.log testing purposes
    // console.log("This is the shirts link: " + findShirtLinks);

    // //call iterateLinks function, pass in the findShirtLinks variable to scrape that page
    // iterateLinks(findShirtLinks);

  }).catch(function(error) {
  console.error(error.message); //end if
  //tell the user in lamens terms why the scrape may have failed.
  console.error('Scrape failed from: ' + homeURL + ' The site may be down, or your connection may need troubleshooting.');
  });//end catch error
 //end scraper

}



//create function to write the CSV file.
function FileWrite() {
  //fields variable holds the column headers
  var fields = ['Title', 'Price', 'ImageURL', 'URL', 'Time'];
  //CSV variable for injecting the fields and object into the converter.
  var csv = json2csv({data: ShirtProps, fields: fields}); 
  console.log(csv);

  //creating a simple date snagger for writing the file with date in the file name.
  var d = new Date();
  var month = d.getMonth()+1;
  var day = d.getDate();
  var output = d.getFullYear() + '-' +
  ((''+month).length<2 ? '0' : '') + month + '-' +
  ((''+day).length<2 ? '0' : '') + day;

  fs.writeFile('./data/' + output + '.csv', csv, function (error) {
          if (error) throw error;
          console.error('There was an error writing the CSV file.');

    });

} //end FileWrite


//Check if data folder exists, source: http://stackoverflow.com/questions/4482686/check-synchronously-if-file-directory-exists-in-node-js
function DataFolderExists(folder) {
  try {
    // Query the entry
    var DataFolder = fs.lstatSync(folder);

    // Is it a directory?
    if (DataFolder.isDirectory()) {
        return true;
    } else {
        return false;
    }
} //end try
catch (error) {
    console.error(error.message);
    console.error('There was an error checking if the folder exists.');
}

}  //end DataFolderExists

【问题讨论】:

标签: jquery node.js npm promise


【解决方案1】:

var elems = $("a[href*=shirt]").nextAll(), var eachLength = elems.length;

使用 nextall() 获取数组中的所有元素。 所以我们现在有了长度,使用这个长度我们可以验证并调用文件写入函数

   'use strict';

    //require NPM packages


    //I chose to use request to make the http calls because it is very easy to use.
    //This npm package also has recent updates, within the last 2 days.
    //Lastly it has a huge number of downloads, this means it has a solid reputation in the community
    var request = require('request');


    //I chose to use cheerio to write the jquery for our node scraper,
    //This package is very simple to use, and it was easy to write jQuery I was already familiar with,
    //Cheerio also makes it simple for us to work with HTML elements on the server.
    //Lastly, Cheerio is popular within the community, with continuous updates and a lot of downloads.
    var cheerio = require('cheerio');

    var rp = require('request-promise');

    var fs = require('fs');


    //I used the json2csv npm package because it was easy to implement into my code,
    //This module also has frequent updates and heavy download activity.
    //This is the most elegant package to download for simple translation of json objects to a CSV file format.
    var json2csv = require('json2csv');




    //Array for shirts JSON object for json2csv to write.
    var ShirtProps = [];
    var Counter = 0;
    var homeURL = "http://www.shirts4mike.com/";


    //start the scraper
    scraper()


    //Initial scrape of the home page, looking for shirts
    function scraper () {

      //use the datafolderexists function to check if data is a directory
      if (!DataFolderExists('data')) {
        fs.mkdir('data');
      }
      //initial request of the home url to find links that may have shirts in them

    rp(homeURL).then(function (html) {

    //use cheerio to load the HTML for scraping
    var $ = cheerio.load(html);
    //For every link with shirt in it iterate over the link and make a request.

    var elems = $("a[href*=shirt]").nextAll(), 
    var eachLength = elems.length;

    elems.each(function() {


        //request promise 
        rp('http://www.shirts4mike.com/' + $(this).attr("href")).then(function (html) {

            //pass the html into the shirt data creator, so if it wound up scraping individual shirts from any of the links it adds it to the data object
            var $ = cheerio.load(html);
            //if the add to cart input exists, log the data to the shirtprops arary.
            if ($('input[value="Add to Cart"]').length) {
              var ShirtURL = $(this).find('a').attr('href');
              var time = new Date();
              //json array for json2csv
              var ShirtData = {
                Title: $('title').html(),
                Price: $('.price').html(),
                ImageURL: $('img').attr('src'),
                URL: homeURL + ShirtURL,
                Time: time.toString() 
              };
                ShirtProps.push(ShirtData);
                console.log(ShirtData);
                Counter ++;
                if (eachLength == Counter ) {
                  FileWrite();
                };
            } else {
              //else we are on a products page, scrape those links for shirt data
                var InnerElm = $('ul.products li').nextAll(), 
                var innereachLength = InnerElm.length;
                var innercount= 0;
                InnerElm.each(function() {
                var ShirtURL = $(this).find('a').attr('href');
                    rp('http://www.shirts4mike.com/' + ShirtURL).then(function (html){
                      innercount++;
                    var $ = cheerio.load(html);
                    var time = new Date();
                    var ShirtData = {
                      Title: $('title').html(),
                      Price: $('.price').html(),
                      ImageURL: $('img').attr('src'),
                      Url: homeURL + ShirtURL,
                      Time: time.toString()
                    };
                    ShirtProps.push(ShirtData);
                    if (innercount == innereachLength) {
                        Counter ++;
                        if (eachLength == Counter ) {
                          FileWrite();
                        };
                    };
                  console.log(ShirtData);

          }).catch(function(error) {
             Counter ++;
            if (eachLength == Counter ) {
                FileWrite();
            };
          console.error(error.message);
          console.error('Scrape failed from: ' + homeURL + 'blah2' + ' The site may be down, or your connection may need troubleshooting.');
          }); //end catch error
      }); //end products li each
              } //end else



    }).catch(function(error) {  //end rp
      console.error(error.message); //end if
  //tell the user in lamens terms why the scrape may have failed.
      console.error('Scrape failed from: ' + homeURL + 'blah' + ' The site may be down, or your connection may need troubleshooting.');
    }); //end catch error
  });  //end href each
    //one thing all shirts links have in common, they are contained in a div with class shirts, find the link to the shirts page based on this class.

    // //console.log testing purposes
    // console.log("This is the shirts link: " + findShirtLinks);

    // //call iterateLinks function, pass in the findShirtLinks variable to scrape that page
    // iterateLinks(findShirtLinks);

  }).catch(function(error) {
  console.error(error.message); //end if
  //tell the user in lamens terms why the scrape may have failed.
  console.error('Scrape failed from: ' + homeURL + ' The site may be down, or your connection may need troubleshooting.');
  });//end catch error
 //end scraper

}



//create function to write the CSV file.
function FileWrite() {
  //fields variable holds the column headers
  var fields = ['Title', 'Price', 'ImageURL', 'URL', 'Time'];
  //CSV variable for injecting the fields and object into the converter.
  var csv = json2csv({data: ShirtProps, fields: fields}); 
  console.log(csv);

  //creating a simple date snagger for writing the file with date in the file name.
  var d = new Date();
  var month = d.getMonth()+1;
  var day = d.getDate();
  var output = d.getFullYear() + '-' +
  ((''+month).length<2 ? '0' : '') + month + '-' +
  ((''+day).length<2 ? '0' : '') + day;

  fs.writeFile('./data/' + output + '.csv', csv, function (error) {
          if (error) throw error;
          console.error('There was an error writing the CSV file.');

    });

} //end FileWrite


//Check if data folder exists, source: http://stackoverflow.com/questions/4482686/check-synchronously-if-file-directory-exists-in-node-js
function DataFolderExists(folder) {
  try {
    // Query the entry
    var DataFolder = fs.lstatSync(folder);

    // Is it a directory?
    if (DataFolder.isDirectory()) {
        return true;
    } else {
        return false;
    }
} //end try
catch (error) {
    console.error(error.message);
    console.error('There was an error checking if the folder exists.');
}

} 

【讨论】:

  • 我喜欢这个想法,但它会多次写入文件,我想在所有请求完成后一次写入文件。
  • 我认为这段代码可以工作。当请求成功或失败时,我们增加计数,当它等于长度时它调用文件写入。
【解决方案2】:
  • 与每个async 操作一样,无论是callbacks 还是promises,在loop 中调用它们时,您应该始终将它们组合在一起。分组方法的选择是您的,但您通常希望使用并行选项。考虑放弃特定的承诺版本的模块,学习更多的general library(无论如何,它通常总是有自己的.promisify()方法)并利用它的.parallel()方法。

  • 在处理nested promises 时,不要忘记始终在return 中包含.then(function(){...} 语句。如果你不这样做,你的 Promise 链将不会知道它必须等待嵌套的 Promise 解决才能继续前进。

  • 您不必为每个 Promise 指定 .catch(function(){...}) 函数,因为错误冒泡的方式与使用常规 try {} catch (e) {} 块代码的方式几乎相同,用于同步操作。

【讨论】:

    【解决方案3】:

    据我所知,request-promise 使用 bluebird 作为承诺。内置了很多辅助方法,请参阅 http://bluebirdjs.com/docs/api-reference.html 了解详情。

    一般:如果你想等待一堆承诺被解决,你可以使用 Promise.all 像:

    var Promise = require("bluebird");
    var promises = [];
    for (var i = 0; i < 100; ++i) {
        promises.push(someAsyncFunction(i));
    }
    Promise.all(promises).then(function() {
        console.log("all the promises were resolved");
    });
    

    ps:在爬虫开始时,您使用异步 fs 方法,但不等待结果。您不想等待 cb 或使用同步之一 (mkdirSync)

    【讨论】:

    • 你能举一个更具体的例子吗?我不明白为什么我们的计数器设置为 100 并且 someasyncfunction 在循环中采用“ i ”的参数,我们不想将函数推送到数组 100 次不?
    • 您在原始示例中没有使用 for 循环吗?我只是假设你想并行请求。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2011-06-27
    • 2016-08-23
    • 1970-01-01
    • 2018-07-10
    • 1970-01-01
    • 2018-02-24
    • 2021-07-25
    相关资源
    最近更新 更多