【问题标题】:R Rvest for() and Error server error: (503) Service UnavailableR Rvest for() 和错误服务器错误:(503)服务不可用
【发布时间】:2015-07-07 21:22:39
【问题描述】:

我是网络抓取的新手,但我很高兴在 R 中使用 rvest。 我试图用它来抓取公司的特定数据。 我创建了一个 for 循环(171 个 url),当我运行它时,它会在第 6 个或第 7 个 url 处停止并出现错误

Error in parse.response(r, parser, encoding = encoding) : 
  server error: (503) Service Unavailable

当我从第 7 个 url 开始循环时,它会再运行两到三个,然后再次停止并出现相同的错误。 我的循环

library(rvest)    
thing<-c("http://www.informazione-aziende.it/Azienda_ LA-VIS-S-C-A",                                                                                  
    "http://www.informazione-aziende.it/Azienda_ L-ANGOLO-DEL-DOLCE-DI-OBEROSLER-MARCO",                                                         
    "http://www.informazione-aziende.it/Azienda_ MARCHI-LAURA",                                                                                 
    "http://www.informazione-aziende.it/Azienda_ LAVIS-PIZZA-DI-GASPARETTO-MATTEO",                                                              
    "http://www.informazione-aziende.it/Azienda_ LE-DELIZIE-MOCHENE-DI-OSLER-NICOLA",                                                            
    "http://www.informazione-aziende.it/Azienda_ LE-DELIZIE-S-N-C-DI-GAMBONI-PIETRO-E-PISONI-MAURO-C-IN-SIGLA-LE-DELIZIE-S-N-C",                 
    "http://www.informazione-aziende.it/Azienda_ LE-FONTI-DISTILLATI-DI-COVI-MARCELLO",                                                          
    "http://www.informazione-aziende.it/Azienda_ LE-MIGOLE-DI-MATTEOTTI-LUCA",                                                                   
    "http://www.informazione-aziende.it/Azienda_ LECHTHALER-DI-TOGN-LUIGI-E-C-S-N-C",                                                            
    "http://www.informazione-aziende.it/Azienda_ LETRARI-AZ-AGRICOLA")

    thing<-gsub(" ", "", thing)

    d <- matrix(nrow=10, ncol=4)
    colnames(d)<-c("RAGIONE SOCIALE",'ATTIVITA', 'INDIRIZZO', 'CAP')

    for(i in 1:10) {
            a<-thing[i]

            urls<-html(a)

            d[i,2] <- try({ urls %>% html_node(".span") %>% html_text() }, silent=TRUE)
    }

可能有办法避免这个错误,提前谢谢你,任何帮助将不胜感激。

UPD 使用下一个代码,我正在尝试从最后一个成功的repeat() 重新开始获取数据的循环,但它正在无限循环,希望有一些建议。

    for(i in 1:10) {

  a<-thing[i]

  try({d[i,2]<- try({html(a) }, silent=TRUE)  %>%
         html_node(".span") %>%
         html_text() }, silent=TRUE)

  repeat {try({d[i,2]<- try({html(a) }, silent=TRUE)  %>%
                 html_node(".span") %>%
                 html_text() }, silent=TRUE)}
  if (!is.na(d[i,2])) break
}

或while()

for(i in 1:10) {

  a<-thing[i]

while (is.na(d[i,2])) {
  try({d[i,2]<-try({html(a) %>%html_node(".span")},silent=TRUE) %>% html_text() },silent=TRUE)
}
}

While() 工作但不是很好而且太慢((

【问题讨论】:

  • 这些网页很可能是not available - 我认为这与rvest 本身没有任何关系。您可以使用 try(..., silent=TRUE) 跳过损坏的 URL。
  • @nrussell ,可能是if else,if 错误else(then) 的某些函数类型,从最后一个+1 结果url 重新启动循环。但我无法想象它的代码)它会有效吗?
  • d[i,2] &lt;- try({ urls %&gt;% html_node(".span") %&gt;% html_text() }, silent=TRUE) 可能会正常工作。
  • @nrussell,谢谢,但仍然平均给出 2 或 3 个,但不多)
  • @DimaSukhorukov,问题是错误发生在html() 调用上,您已将其置于try() 块之外。将urls&lt;-html(a) 语句移到try() 块内,它会起作用。

标签: r loops error-handling scrape rvest


【解决方案1】:

看起来如果您访问该网站的速度太快,您会得到 503。添加一个 Sys.sleep(2) 并且所有 10 次迭代都对我有用...

library(rvest)    
thing<-c("http://www.informazione-aziende.it/Azienda_ LA-VIS-S-C-A",                                                                                  
         "http://www.informazione-aziende.it/Azienda_ L-ANGOLO-DEL-DOLCE-DI-OBEROSLER-MARCO",                                                         
         "http://www.informazione-aziende.it/Azienda_ MARCHI-LAURA",                                                                                 
         "http://www.informazione-aziende.it/Azienda_ LAVIS-PIZZA-DI-GASPARETTO-MATTEO",                                                              
         "http://www.informazione-aziende.it/Azienda_ LE-DELIZIE-MOCHENE-DI-OSLER-NICOLA",                                                            
         "http://www.informazione-aziende.it/Azienda_ LE-DELIZIE-S-N-C-DI-GAMBONI-PIETRO-E-PISONI-MAURO-C-IN-SIGLA-LE-DELIZIE-S-N-C",                 
         "http://www.informazione-aziende.it/Azienda_ LE-FONTI-DISTILLATI-DI-COVI-MARCELLO",                                                          
         "http://www.informazione-aziende.it/Azienda_ LE-MIGOLE-DI-MATTEOTTI-LUCA",                                                                   
         "http://www.informazione-aziende.it/Azienda_ LECHTHALER-DI-TOGN-LUIGI-E-C-S-N-C",                                                            
         "http://www.informazione-aziende.it/Azienda_ LETRARI-AZ-AGRICOLA")

thing<-gsub(" ", "", thing)

d <- matrix(nrow=10, ncol=4)
colnames(d)<-c("RAGIONE SOCIALE",'ATTIVITA', 'INDIRIZZO', 'CAP')

for(i in 1:10) {
  print(i)
  a<-thing[i]  
  urls<-html(a)  
  d[i,2] <- try({ urls %>% html_node(".span") %>% html_text() }, silent=TRUE)
  Sys.sleep(2)
}

【讨论】:

  • 哇,这真的很酷!!,Sys.sleep(2) 休息 2 秒吗??
  • 是的,不确定 2 秒是足够还是太多。此外,合并其他一些解决方案来清理代码可能是明智的......
猜你喜欢
  • 2017-01-14
  • 2012-04-23
  • 1970-01-01
  • 2011-08-04
  • 1970-01-01
  • 2021-01-07
  • 1970-01-01
  • 2015-02-28
相关资源
最近更新 更多