【问题标题】:“Null” error with some tables on same page with R readHTMLTable使用 R readHTMLTable 在同一页面上的某些表出现“空”错误
【发布时间】:2018-08-28 01:22:28
【问题描述】:

我正在尝试在下一页上抓取表格数据 HKJC LINK

我可以读取表2和表3,但是表1返回“NULL”,我检查了包含数据(see the picture)的源代码

    library(XML)
    url = "http://racing.hkjc.com/racing/info/meeting/Results/english/Local/20180715/ST/6"
    sample=readHTMLTable(url,which=2,encoding = "UTF-8")
    head(sample,1)

有什么想法吗?

【问题讨论】:

    标签: r web-scraping


    【解决方案1】:
    library(XML)
    
    scrapescrape <- function(x) {
    
      link <- paste0("http://www.racingpost.com/horses/horse_home.sd?horse_id=",x)
    
        tryCatch(readHTMLTable(link, which=2), error=function(e){NA})
    
      }
    }
    
    ids <- c(896119, 766254, 790946, 556341,  62736, 660506, 486791, 580134, 0011, 580134)
    
    lst <- lapply(ids, scrapescrape)
    
    str(lst)
    

    【讨论】:

      猜你喜欢
      • 2013-08-04
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-09-16
      • 1970-01-01
      • 2023-03-27
      • 2012-08-25
      相关资源
      最近更新 更多