【问题标题】:Add unique identifier column to a nested list of data frames将唯一标识符列添加到数据框的嵌套列表中
【发布时间】:2021-11-30 14:09:11
【问题描述】:

示例数据

我正在处理数据框的嵌套列表。我的列表包含 1,000 多个列表,每个列表都包含一个数据框作为其唯一元素。每个数据框包含 10 多个变量的 30 多个观察值。为简单起见,这里有一个小示例列表:

df1 <- tibble::tibble(a = 1:25, b = 1:25, c = 1:25, d = 1:25, e = 1:25)
df2 <- tibble::tibble(a = 1:35, b = 1:35, c = 1:35, d = 1:35, e = 1:35)
df3 <- tibble::tibble(a = 1:30, b = 1:30, c = 1:30, d = 1:30, e = 1:30)
df4 <- tibble::tibble(a = 1:20, b = 1:20, c = 1:20, d = 1:20, e = 1:20)

dfs_list <- list(list(a = df1), list(a = df2), list(a = df3), list(a = df4))
  dfs_list

[[1]]
[[1]][[1]]
# A tibble: 25 x 5
       a     b     c     d     e
   <int> <int> <int> <int> <int>
 1     1     1     1     1     1
 2     2     2     2     2     2
 3     3     3     3     3     3
 4     4     4     4     4     4
 5     5     5     5     5     5
 6     6     6     6     6     6
 7     7     7     7     7     7
 8     8     8     8     8     8
 9     9     9     9     9     9
10    10    10    10    10    10
# ... with 15 more rows


[[2]]
[[2]][[1]]
# A tibble: 35 x 5
       a     b     c     d     e
   <int> <int> <int> <int> <int>
 1     1     1     1     1     1
 2     2     2     2     2     2
 3     3     3     3     3     3
 4     4     4     4     4     4
 5     5     5     5     5     5
 6     6     6     6     6     6
 7     7     7     7     7     7
 8     8     8     8     8     8
 9     9     9     9     9     9
10    10    10    10    10    10
# ... with 25 more rows


[[3]]
[[3]][[1]]
# A tibble: 30 x 5
       a     b     c     d     e
   <int> <int> <int> <int> <int>
 1     1     1     1     1     1
 2     2     2     2     2     2
 3     3     3     3     3     3
 4     4     4     4     4     4
 5     5     5     5     5     5
 6     6     6     6     6     6
 7     7     7     7     7     7
 8     8     8     8     8     8
 9     9     9     9     9     9
10    10    10    10    10    10
# ... with 20 more rows


[[4]]
[[4]][[1]]
# A tibble: 20 x 5
       a     b     c     d     e
   <int> <int> <int> <int> <int>
 1     1     1     1     1     1
 2     2     2     2     2     2
 3     3     3     3     3     3
 4     4     4     4     4     4
 5     5     5     5     5     5
 6     6     6     6     6     6
 7     7     7     7     7     7
 8     8     8     8     8     8
 9     9     9     9     9     9
10    10    10    10    10    10
# ... with 10 more rows

所需的输出

我正在尝试为列表中的每个数据框生成一个包含唯一标识符的列。该列将基于两个数字序列;比如说1:10 和1:100。例如,第一个数据框中的列将包含 1.1,第二个将包含 2.1,依此类推,一直到 10.100.

从头开始处理较小的示例,让我们制作我的数字序列1:2 和1:2。下面的identifier 列是我希望添加到列表中每个数据框的内容:

[[1]]
[[1]][[1]]
# A tibble: 25 x 6
       a     b     c     d     e identifier
   <int> <int> <int> <int> <int> <chr>     
 1     1     1     1     1     1 1.1       
 2     2     2     2     2     2 1.1       
 3     3     3     3     3     3 1.1       
 4     4     4     4     4     4 1.1       
 5     5     5     5     5     5 1.1       
 6     6     6     6     6     6 1.1       
 7     7     7     7     7     7 1.1       
 8     8     8     8     8     8 1.1       
 9     9     9     9     9     9 1.1       
10    10    10    10    10    10 1.1       
# ... with 15 more rows


[[2]]
[[2]][[1]]
# A tibble: 35 x 6
       a     b     c     d     e identifier
   <int> <int> <int> <int> <int> <chr>     
 1     1     1     1     1     1 2.1      
 2     2     2     2     2     2 2.1       
 3     3     3     3     3     3 2.1       
 4     4     4     4     4     4 2.1       
 5     5     5     5     5     5 2.1       
 6     6     6     6     6     6 2.1       
 7     7     7     7     7     7 2.1       
 8     8     8     8     8     8 2.1       
 9     9     9     9     9     9 2.1       
10    10    10    10    10    10 2.1       
# ... with 25 more rows


[[3]]
[[3]][[1]]
# A tibble: 30 x 6
       a     b     c     d     e identifier
   <int> <int> <int> <int> <int> <chr>     
 1     1     1     1     1     1 1.2       
 2     2     2     2     2     2 1.2       
 3     3     3     3     3     3 1.2       
 4     4     4     4     4     4 1.2       
 5     5     5     5     5     5 1.2       
 6     6     6     6     6     6 1.2       
 7     7     7     7     7     7 1.2       
 8     8     8     8     8     8 1.2       
 9     9     9     9     9     9 1.2       
10    10    10    10    10    10 1.2       
# ... with 20 more rows


[[4]]
[[4]][[1]]
# A tibble: 20 x 5
       a     b     c     d     e identifier
   <int> <int> <int> <int> <int> <chr>     
 1     1     1     1     1     1 2.2       
 2     2     2     2     2     2 2.2       
 3     3     3     3     3     3 2.2       
 4     4     4     4     4     4 2.2       
 5     5     5     5     5     5 2.2       
 6     6     6     6     6     6 2.2       
 7     7     7     7     7     7 2.2       
 8     8     8     8     8     8 2.2       
 9     9     9     9     9     9 2.2       
10    10    10    10    10    10 2.2       
# ... with 10 more rows

尝试的方法

我尝试使用apply(expand.grid()) 创建一个数组,然后使用mapply() 将数组的一个观察值绑定到每个数据帧:

a.b <- apply(expand.grid(c(1:2), c(1:2)), 1, paste, collapse = '.')
mapply(cbind, dfs_list, "Identifier" = a.b, SIMPLIFY = F)

但是,该列被插入到父列表中,而不是直接插入到数据框中:

[[1]]
               Identifier    
[1,] tbl_df,5 "1.1"

[[2]]
               Identifier    
[1,] tbl_df,5 "2.1"

[[3]]
               Identifier    
[1,] tbl_df,5 "1.2"

[[4]]
               Identifier    
[1,] tbl_df,5 "2.2"

经过反复试验,我在晚上晚些时候尝试了一种稍微不同的方法。起初我以为我已经解决了我的问题,但是生成的列表是 13 GB 而不是之前的 19 MB,并且(相对)花费了(相对)更长的时间来编写,所以我怀疑这是否是解决方案。今天早上我也无法使用我的示例数据集重现我的结果。

> dfs_identify <- dfs_list %>% 
+   apply(function(z) mapply(cbind, z, "Identifier" = a.b, SIMPLIFY = F))
Error in match.fun(FUN) : argument "FUN" is missing, with no default

【问题讨论】:

    标签: r dataframe nested-lists uniqueidentifier mapply


    【解决方案1】:

    您可以使用lapply 和Map 尝试这种方法-

    result <- lapply(seq_along(dfs_list), function(x) {
      Map(cbind, dfs_list[[x]], 
                 Identifier = paste(x, seq_along(dfs_list[[x]]), sep = '.'))
    })
    
    result
    
    #[[1]]
    #[[1]][[1]]
    #                   mpg cyl disp  hp drat Identifier
    #Mazda RX4         21.0   6  160 110 3.90        1.1
    #Mazda RX4 Wag     21.0   6  160 110 3.90        1.1
    #Datsun 710        22.8   4  108  93 3.85        1.1
    #Hornet 4 Drive    21.4   6  258 110 3.08        1.1
    #Hornet Sportabout 18.7   8  360 175 3.15        1.1
    
    #[[1]][[2]]
    #                   mpg cyl disp  hp drat Identifier
    #Mazda RX4 Wag     21.0   6  160 110 3.90        1.2
    #Datsun 710        22.8   4  108  93 3.85        1.2
    #Hornet 4 Drive    21.4   6  258 110 3.08        1.2
    #Hornet Sportabout 18.7   8  360 175 3.15        1.2
    
    
    #[[2]]
    #[[2]][[1]]
    #                   mpg cyl disp  hp drat Identifier
    #Mazda RX4         21.0   6  160 110 3.90        2.1
    #Mazda RX4 Wag     21.0   6  160 110 3.90        2.1
    #Datsun 710        22.8   4  108  93 3.85        2.1
    #Hornet 4 Drive    21.4   6  258 110 3.08        2.1
    #Hornet Sportabout 18.7   8  360 175 3.15        2.1
    

    数据

    如果您在reproducible format 中提供数据会更容易提供帮助

    dfs_list <- list(list(mtcars[1:5, 1:5],mtcars[2:5, 1:5]), list(mtcars[1:5, 1:5]))
    

    【讨论】:

    • 我已经编辑了我的帖子以包含可重复的数据,谢谢!这是解决这个问题的一个强有力的解决方案,我相信我将来会使用它。但是,在我的特定场景中,生成的标识符并没有按照我想要的方式进行迭代——这是因为我的每个嵌套列表都只包含一个数据框。当我使用这种方法时,只有第一个数字按顺序继续,即1:1、2:1、3:1、4:1 等,但我希望第二个数字也能改变,独立/不考虑列表位置:例如,1:1、2:1、1:2、2:2。
    【解决方案2】:
    Map(`names<-`, dfs_list, a.b)
    

    这会为每个列表项指定您创建的名称。它没有说“标识符”,但我认为这就是你所追求的。

    编辑:

    Map(function(x, y) list(cbind(x[[1]], "Identifier" = y)), dfs_list, a.b)
    

    这提供了一个新的标识符列。 x[[1]] 是进入嵌套结构内部,list() 重新创建原始嵌套结构。 map 同 mapply(..., simple = FALSE)

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-03-08
      • 2021-01-15
      相关资源
      最近更新 更多