【发布时间】:2020-12-18 01:32:08
【问题描述】:
我正在尝试从 footballdb.com 进行网络抓取,以获取与我正在通过以下链接创建的模型的 NFL 球员受伤相关的数据:https://www.footballdb.com/transactions/injuries.html?yr=2016&wk=1&type=reg。页面上的所有表格都存储在 divtable 元素中,我可以访问这些元素,但是我无法从每个 divtable 中提取我需要的单个元素(即球员姓名、伤害、wed_status、thurs_status、fri_status、game_status )。有没有人有从 divtable 类中提取元素的经验?以下是我到目前为止的代码。
#Set the start web url
path1<-("https://www.footballdb.com/transactions/injuries.html?yr=")
seasons<-c("2016", "2017", "2020")
weeks<-1:17
data<-NULL
for (i in 1:length(seasons)) {
path2<-paste0(path1,seasons[i])
for (j in 1:length(weeks)) {
path3<-paste0(path2,"&wk=",j,"&type=reg")
URL<-read_html(path3)
divtables<-html_nodes(URL, ".divtable")
for (k in 1:length(divtables)) {
}
}
}
【问题讨论】:
标签: r web-scraping rvest