【问题标题】:Struggling to extract elements of divtable type from webpage using rvest努力使用 rvest 从网页中提取 divtable 类型的元素
【发布时间】:2020-12-18 01:32:08
【问题描述】:

我正在尝试从 footballdb.com 进行网络抓取,以获取与我正在通过以下链接创建的模型的 NFL 球员受伤相关的数据:https://www.footballdb.com/transactions/injuries.html?yr=2016&wk=1&type=reg。页面上的所有表格都存储在 divtable 元素中,我可以访问这些元素,但是我无法从每个 divtable 中提取我需要的单个元素(即球员姓名、伤害、wed_status、thurs_status、fri_status、game_status )。有没有人有从 divtable 类中提取元素的经验?以下是我到目前为止的代码。

#Set the start web url
path1<-("https://www.footballdb.com/transactions/injuries.html?yr=")


seasons<-c("2016", "2017", "2020")
weeks<-1:17
data<-NULL
for (i in 1:length(seasons)) {
  path2<-paste0(path1,seasons[i])
  
  for (j in 1:length(weeks)) {
    path3<-paste0(path2,"&wk=",j,"&type=reg") 
    URL<-read_html(path3)
    divtables<-html_nodes(URL, ".divtable")
    
    for (k in 1:length(divtables)) {
      
    }
  }
}

【问题讨论】:

    标签: r web-scraping rvest


    【解决方案1】:

    我确信有一种更好的方法可以生成所需的 url,但是使用循环格式,您可以生成 url,然后使用 map_dfr 将每个页面的数据帧连接到一个总体 df 中。您可以使用类名获取每个“列”。 “表格”具有代表tr(表格行)和td(表格单元格)的类,例如.tr 而不是 tr,允许您选择要分配到数据框中的列。见css class selectors versus type。我从@RonakShah here 获取了 map_dfr 想法中的数据框。

    library(purrr)
    library(rvest)
    
    path1 <- 'https://www.footballdb.com/transactions/injuries.html?yr='
    seasons <- c("2016", "2017", "2020")
    weeks <- 1:17
    results <- list()
    c <- 1
    
    for(s in seq_along(seasons)){
      for(w in seq_along(weeks)){
        c <- c+1
        results[c] <- paste0(path1, seasons[s],"&wk=", as.character(w), "&type=reg")
      }
    }
    
    #use `results` for all rather than just `results[1:3]`
    
    result <- map_df(compact(results[1:3]), function(x){ 
      page <- read_html(x)
      data.frame(
        Player = page %>% html_nodes('.divtable .td:nth-child(1) b') %>% html_text(),
        Injury = page %>% html_nodes('.divtable .td:nth-child(2)') %>% html_text(),
        Wed = page %>% html_nodes('.divtable .td:nth-child(3)') %>% html_text(),
        Thu = page %>% html_nodes('.divtable .td:nth-child(4)') %>% html_text(),
        Fri = page %>% html_nodes('.divtable .td:nth-child(5)') %>% html_text(),
        GameStatus = page %>% html_nodes('.divtable .td:nth-child(6)') %>% html_text()
      )
      }
    )
    

    【讨论】:

      猜你喜欢
      • 2021-04-19
      • 2021-08-15
      • 1970-01-01
      • 2023-03-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多