【问题标题】:Merge rows when twocolumns have matches in R当两列在R中匹配时合并行
【发布时间】:2021-03-04 09:49:12
【问题描述】:

我有一个数据框,例如 ;

Species   Family  Events  Groups
Monkey    A       6,7     G1,G2
Monkey    A,B     6,8,9   G1,G2,G4,G8,G12
Elephant  B       7,8     G6,G7
Elephant  C       9,10    G6
Dog       K       10      G90
Dog       L,M,N   8,10,9  G90,G91 

这个想法是在Species 内合并,至少在Events 和Groups 列之间存在匹配的列。

例如Monkey:

Species   Family  Events  Groups
Monkey    A       6,7     G1,G2
Monkey    A,B     6,8,9   G1,G2,G4,G8,G12

Event 6 和 row1 中的Groups G1 也在 *row2 中,所以我将它们合并:

Species   Family  Events  Groups
Monkey    A,B       6,7,8,9 G1,G2,G4,G8,G12

最后的预期输出是:

Species   Family  Events  Groups
Monkey    A,B     6,7,8,9 G1,G2,G4,G8,G12
Elephant  B       7,8     G6,G7
Elephant  C       9,10    G6
Dog       K,L,M,N   8,9,10  G90,G91

我没有合并 Elephant,因为 Events 列中没有匹配项。

有人知道代码吗,谢谢。

这是数据:

structure(list(Species = structure(c(3L, 3L, 2L, 2L, 1L, 1L), .Label = c("Dog", 
"Elephant", "Monkey"), class = "factor"), Family = structure(1:6, .Label = c("A", 
"A,B", "B", "C", "K", "L,M,N"), class = "factor"), Events = structure(c(2L, 
3L, 4L, 6L, 1L, 5L), .Label = c("10", "6,7", "6,8,9", "7,8", 
"8,10,9", "9,10"), class = "factor"), Groups = structure(c(1L, 
3L, 2L, 4L, 5L, 6L), .Label = c(" G1,G2", " G6,G7", "G1,G2,G4,G8,G12", 
"G6", "G90", "G90,G91"), class = "factor")), class = "data.frame", row.names = c(NA, 
-6L))

【问题讨论】:

  • 1) 分离,2) pivot_longer 3) 过滤那些 max(n()) >1, 4) 将这些记录串起来,5) 将这些与未过滤的记录合并。
  • 弄错了对不起

标签: r dataframe dplyr


【解决方案1】:

遵循这个策略

library(tidyverse)

df1 <- df %>% 
  group_by(Species) %>% 
  mutate(across(c(Family, Events, Groups), ~as.character(.))) %>%
  summarise(across(c(Events, Groups), ~ toString(Reduce(intersect, strsplit(., ','))))) %>%
  filter(Events != "" & Groups != "") %>%
  select(Species) 

df1 %>%
  left_join(df %>% mutate(across(c(Family, Events, Groups), ~as.character(.)))) %>%
  group_by(Species) %>%
  summarise(across(c(Family, Events, Groups), ~ toString(Reduce(union, strsplit(., ','))))) %>%
  rbind(df %>% anti_join(df1))

# A tibble: 4 x 4
  Species  Family     Events     Groups             
  <fct>    <chr>      <chr>      <chr>              
1 Dog      K, L, M, N 10, 8, 9   G90, G91           
2 Monkey   A, B       6, 7, 8, 9 G1, G2, G4, G8, G12
3 Elephant B          7,8        G6,G7              
4 Elephant C          9,10       G6

【讨论】:

  • 哇,非常感谢您和您的时间!!!它有很大帮助。我也有类似的问题,也许你也可以帮忙??如果你有时间,这里是帖子:stackoverflow.com/questions/66475914/…
猜你喜欢
  • 1970-01-01
  • 2021-12-23
  • 1970-01-01
  • 2021-08-15
  • 2018-12-08
  • 1970-01-01
  • 1970-01-01
  • 2015-04-24
  • 2016-10-19
相关资源
最近更新 更多