【问题标题】:Scraping with rvest: Getting error HTTP 502使用 rvest 抓取:出现错误 HTTP 502
【发布时间】:2019-11-14 15:57:53
【问题描述】:

我有一个 R 脚本,它使用 rvest 从 accuweather 中提取一些数据。 accuweather URL 中包含与城市唯一对应的 ID。我正在尝试提取给定范围内的 ID 和相关的城市名称。 rvest 本身非常适合单个 ID,但是当我遍历 for 循环时,它最终会返回此错误 - “Open.connection(x, "rb") 中的错误:HTTP 错误 502。”

我怀疑这个错误是由于网站阻止了我。我该如何解决这个问题?我想从相当大的范围(10,000 个 ID)中抓取,并且在循环约 500 次迭代后它一直给我这个错误。我也尝试了closeAllConnections()Sys.sleep() 但无济于事。我真的很感激这个问题的任何帮助。

编辑:已解决。我在这里通过这个线程找到了解决方法:Use tryCatch skip to next value of loop upon error?。我使用tryCatch()error = function(e) e 作为参数,它抑制了错误消息并允许循环继续而不会中断。希望这对遇到类似问题的其他人有所帮助。

library(rvest)
library(httr)

# create matrix to store IDs and Cities
# each ID corresponds to a single city 
id_mat<- matrix(0, ncol = 2, nrow = 10001 )

# initialize index for matrix row  
j = 1

for (i in 300000:310000){
  z <- as.character(i)
# pull city name from website 
  accu <- read_html(paste("https://www.accuweather.com/en/us/new-york-ny/10007/june-weather/", z, sep = ""))
  citystate <- accu %>% html_nodes('h1') %>% html_text()
# store values
  id_mat[j,1] = i
  id_mat[j,2] = citystate
# increment by 1 
  i = i + 1 
  j = j + 1
    # close connection after 200 pulls, wait 5 mins and loop again
    if (i %% 200 == 0) {
        closeAllConnections()
        Sys.sleep(300)
        next 
  } else {
        # sleep for 1 or 2 seconds every loop
        Sys.sleep(sample(2,1))
  }
}

【问题讨论】:

    标签: web-scraping rvest http-error


    【解决方案1】:

    问题似乎来自科学记数法。

    How to disable scientific notation?

    我稍微更改了您的代码,现在它似乎可以工作了:

    library(rvest)
    library(httr)
    
    id_mat<- matrix(0, ncol = 2, nrow = 10001 )
    
    readUrl <- function(url) {
    out <- tryCatch(
    {   
      download.file(url, destfile = "scrapedpage.html", quiet=TRUE)
      return(1)
    },
    error=function(cond) {
    
      return(0)
    },
    warning=function(cond) {
      return(0)
    }
    )    
    return(out)
    }
    
    j = 1
    
    options(scipen = 999)
    
    for (i in 300000:310000){
      z <- as.character(i)
    # pull city name from website 
      url <- paste("https://www.accuweather.com/en/us/new-york-ny/10007/june-weather/", z, sep = "")
      if( readUrl(url)==1) {
      download.file(url, destfile = "scrapedpage.html", quiet=TRUE)
      accu <- read_html("scrapedpage.html")
      citystate <- accu %>% html_nodes('h1') %>% html_text()
    # store values
      id_mat[j,1] = i
      id_mat[j,2] = citystate
    # increment by 1 
      i = i + 1 
      j = j + 1
        # close connection after 200 pulls, wait 5 mins and loop again
        if (i %% 200 == 0) {
            closeAllConnections()
            Sys.sleep(300)
            next 
      } else {
            # sleep for 1 or 2 seconds every loop
            Sys.sleep(sample(2,1))
      }
       } else {er <- 1}
      }
    

    【讨论】:

    猜你喜欢
    • 2015-09-14
    • 1970-01-01
    • 1970-01-01
    • 2023-01-07
    • 1970-01-01
    • 2017-03-05
    • 1970-01-01
    • 2018-12-01
    • 2013-09-17
    相关资源
    最近更新 更多