【问题标题】:Fast way to parse vector of "continent / country / city" in R在R中解析“大陆/国家/城市”向量的快速方法
【发布时间】:2021-09-15 23:33:53
【问题描述】:

我在 R 中有一个字符向量,每个字符串由“大陆/国家/城市”组成,例如

x=rep("Africa / Kenya / Nairobi", 1000000)

但“/”有时会在没有括号空格的情况下被错误输入为“/”,并且在某些情况下城市也会丢失,因此它会例如是“非洲/肯尼亚”,没有城市。

我想将其解析为三个向量大陆、国家和城市,如果缺少城市,则使用 NA。

对于国家我现在做了类似的事情

country = sapply(x, function(loc) trimws(strsplit(loc,"/", fixed = TRUE)[[1]][2]))

但是如果向量 x 很长,那会很慢。在 R 中解析这个的有效方法是什么?

【问题讨论】:

  • strsplit 已经被矢量化了,所以最好直接调用它而不是在那里使用sapply。但是“非常慢”的确切定义是什么,“更高效”的结果的要求是什么?如果你想编写自己的解析器,如果性能是一个非常重要的问题,你总是可以使用 Rcpp 编写自己的 C++ 代码。

标签: r text-parsing strsplit


【解决方案1】:

您可以在do.call 中尝试rbind。在lapply 中使用[ 是为了得到3 个结果,以防城市丢失。

x <- c("Africa / Kenya / Nairobi", "Africa/Kenya/Nairobi", "Africa / Kenya")

y <- do.call(rbind, lapply(strsplit(x, "/", TRUE), "[", 1:3))
y <- trimws(y, whitespace = " ")

y
#     [,1]     [,2]    [,3]     
#[1,] "Africa" "Kenya" "Nairobi"
#[2,] "Africa" "Kenya" "Nairobi"
#[3,] "Africa" "Kenya" NA       

或者使用data.table:

x <- c("Africa / Kenya / Nairobi", "Africa/Kenya/Nairobi", "Africa / Kenya")

y <- do.call(cbind, data.table::tstrsplit(x, "/", TRUE))
y <- trimws(y, whitespace = " ")

y
#     [,1]     [,2]    [,3]     
#[1,] "Africa" "Kenya" "Nairobi"
#[2,] "Africa" "Kenya" "Nairobi"
#[3,] "Africa" "Kenya" NA       

基准测试

#x <- rep("Africa / Kenya / Nairobi", 1000000) #Timings will depend on the used dataset

n <- 1e6L
f1 <- function(n) replicate(n, paste(sample(letters, sample(5:15, 1), TRUE), collapse = ""))
f2 <- function(n) sample(c("/", " /", "/ ", " / "), n, TRUE)
set.seed(42)
x <- paste0(f1(n), f2(n), f1(n), sample(c(paste0(f2(n%/%2L), f1(n%/%2L)), rep("", n - n%/%2L))))

system.time( #Method given in the question
  sapply(x, function(loc) trimws(strsplit(loc,"/", fixed = TRUE)[[1]][2])))
#       User      System verstrichen 
#     47.718       0.004      47.798 

system.time(  #Using strsplit and trimws
  trimws(do.call(rbind, lapply(strsplit(x, "/", TRUE), "[", 1:3)), whitespace = " "))
#       User      System verstrichen 
#      5.446       0.008       5.454 

system.time(  #Using data.table::tstrsplit and trimws
  trimws(do.call(cbind, data.table::tstrsplit(x, "/", TRUE)), whitespace = " "))
#       User      System verstrichen 
#      2.365       0.012       2.376 

system.time(  #Using readr::read_delim from @Anoushiravan R
  readr::read_delim(x, delim = "/", quote = "", trim_ws = TRUE, col_names = FALSE))
#       User      System verstrichen 
#      1.961       0.024       2.222 

system.time(  #Using data.table::tstrsplit with " */ *"
  do.call(cbind, data.table::tstrsplit(x, " */ *", perl=TRUE)))
#       User      System verstrichen 
#      1.394       0.000       1.394 

system.time(  #Using read.table from @akrun
  read.table(text = x, sep = "/", header = FALSE, fill = TRUE, strip.white = TRUE, na.strings = ""))
#       User      System verstrichen 
#      1.298       0.004       1.302 

system.time(  #Using data.table::fread from @akrun
  data.table::fread(text = paste(x, collapse="\n"), sep="/", fill = TRUE, na.strings = ""))
#       User      System verstrichen 
#      1.146       0.016       0.996 

system.time(  #Using read.table with additional argiments
  read.table(text = x, sep = "/", header = FALSE, fill = TRUE, strip.white = TRUE, na.strings = "", nrows=length(x), comment.char = "", colClasses = c("character")))
#       User      System verstrichen 
#      1.076       0.000       1.076 

system.time(  #Using data.table::fread with stringr::str_c (or stringi::stri_c)
  data.table::fread(text = stringr::str_c(x, collapse="\n"), sep="/", fill = TRUE, na.strings = ""))
#       User      System verstrichen 
#      0.780       0.000       0.624 

使用data.table::fread 并使用stringr::str_c 创建输入字符串看起来是当前给定方法中最快的。

【讨论】:

  • 非常感谢 - 这正是我想要的!
【解决方案2】:

考虑从base R 使用read.table

read.table(text = x, sep = "/", header = FALSE,
      fill = TRUE, strip.white = TRUE, na.strings = "")
      V1    V2      V3
1 Africa Kenya Nairobi
2 Africa Kenya Nairobi
3 Africa Kenya    <NA>

或者使用来自data.table的fread

library(data.table)
fread(text = paste(x, collapse="\n"), sep="/", fill = TRUE, na.strings = "")
   Africa Kenya Nairobi
1: Africa Kenya Nairobi
2: Africa Kenya    <NA>

基准测试

x <- rep("Africa / Kenya / Nairobi", 1000000)
> 
> system.time(fread(text = paste(x, collapse="\n"), sep="/", fill = TRUE, na.strings = ""))
   user  system elapsed 
  0.473   0.024   0.496 

> system.time(read.table(text = x, sep = "/", header = FALSE,
+       fill = TRUE, strip.white = TRUE, na.strings = ""))
   user  system elapsed 
  0.519   0.026   0.543 

> system.time({  #Using data.table
+   y <- do.call(cbind, data.table::tstrsplit(x, "/", TRUE))
+   y <- trimws(y, whitespace = " ")
+ })
   user  system elapsed 
  2.035   0.051   2.067 

数据

x <- c("Africa / Kenya / Nairobi", "Africa/Kenya/Nairobi", "Africa / Kenya")

【讨论】:

  • 这也很优雅...这比上面的 data.table 解决方案慢还是快?我猜慢一点,对吧?
  • @TomWenseleers 请检查基准。这两个选项似乎都比其他帖子中显示的最快速度更快
  • read.table 总是为这类问题发挥魔力,干杯!
  • 非常感谢!只是没有考虑过这个选项,即使我一直使用 read.table 。 :-) 鉴于这是最快的选项,我已将其作为正确答案进行了检查,尽管我发现另一个选项也非常具有指导意义......
【解决方案3】:

我觉得这个也可以用:

library(readr)

xx <- readr::read_delim(b, delim = "/", quote = "", trim_ws = TRUE, col_names = FALSE)

# A tibble: 3 x 3
  X1     X2    X3     
  <chr>  <chr> <chr>  
1 Africa Kenya Nairobi
2 Africa Kenya Nairobi
3 Africa Kenya NA   

【讨论】:

  • 非常感谢亲爱的阿伦,我不知道。
  • 另外,col_names = FALSE 又返回一行
  • 哦,否则它会使用第一行作为colnames,有趣。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2022-06-15
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-08-22
  • 2017-10-20
  • 2011-01-14
相关资源
最近更新 更多