【发布时间】:2021-03-04 09:49:12
【问题描述】:
我有一个数据框,例如 ;
Species Family Events Groups
Monkey A 6,7 G1,G2
Monkey A,B 6,8,9 G1,G2,G4,G8,G12
Elephant B 7,8 G6,G7
Elephant C 9,10 G6
Dog K 10 G90
Dog L,M,N 8,10,9 G90,G91
这个想法是在Species 内合并,至少在Events 和Groups 列之间存在匹配的列。
例如Monkey:
Species Family Events Groups
Monkey A 6,7 G1,G2
Monkey A,B 6,8,9 G1,G2,G4,G8,G12
Event 6 和 row1 中的Groups G1 也在 *row2 中,所以我将它们合并:
Species Family Events Groups
Monkey A,B 6,7,8,9 G1,G2,G4,G8,G12
最后的预期输出是:
Species Family Events Groups
Monkey A,B 6,7,8,9 G1,G2,G4,G8,G12
Elephant B 7,8 G6,G7
Elephant C 9,10 G6
Dog K,L,M,N 8,9,10 G90,G91
我没有合并 Elephant,因为 Events 列中没有匹配项。
有人知道代码吗,谢谢。
这是数据:
structure(list(Species = structure(c(3L, 3L, 2L, 2L, 1L, 1L), .Label = c("Dog",
"Elephant", "Monkey"), class = "factor"), Family = structure(1:6, .Label = c("A",
"A,B", "B", "C", "K", "L,M,N"), class = "factor"), Events = structure(c(2L,
3L, 4L, 6L, 1L, 5L), .Label = c("10", "6,7", "6,8,9", "7,8",
"8,10,9", "9,10"), class = "factor"), Groups = structure(c(1L,
3L, 2L, 4L, 5L, 6L), .Label = c(" G1,G2", " G6,G7", "G1,G2,G4,G8,G12",
"G6", "G90", "G90,G91"), class = "factor")), class = "data.frame", row.names = c(NA,
-6L))
【问题讨论】:
-
1) 分离,2) pivot_longer 3) 过滤那些 max(n()) >1, 4) 将这些记录串起来,5) 将这些与未过滤的记录合并。
-
弄错了对不起