【问题标题】:How to create a new column from existing two columns but omitting NAs rows in R如何从现有的两列创建一个新列但在 R 中省略 NAs 行
【发布时间】:2023-01-18 01:18:51
【问题描述】:

我有一个数据框,它的一部分看起来像这样:

Domain <- c(rep("Bacteria",3),rep("Archaea", 2))
Phylum <- c("Proteobacteria","Cyanobacteria","Planctomycetota", "Thermoplasmatota", "Thermoplasmatota")
Class <- c("Alphaproteobacteria","Cyanobacteriia","Phycisphaerae","Poseidoniia_A",NA)
Order <- c("Sphingomonadales", NA, "Phycisphaerales", "Poseidoniales", NA)
Family <- c("Emcibacteraceae", NA, NA, "Poseidonia", NA)
Genus <- c("UBA4441", NA,NA,NA,NA)
Species <- c("UBA4441 sp", NA,NA,NA,NA)


demo_table <- data.frame(Domain, Phylum, Class, Order, Family, Genus, Species)

这里的要点是我想创建一个名为“分配”的新列,该列包含逐行包含非 NA 值的最后两列的合并,并且这些值由空格分隔。

这是预期的输出:

Domain Phylum Class Order Family Genus Species assignation
Bacteria Proteobacteria Alphaproteobacteria Sphingomonadales Emcibacteraceae UBA4441 UBA4441 sp UBA4441 UBA4441 sp
Bacteria Cyanobacteria Cyanobacteriia NA NA NA NA Cyanobacteria Cyanobacteriia
Bacteria Planctomycetota Phycisphaerae Phycisphaerales NA NA NA Phycisphaerae Phycisphaerales
Archaea Thermoplasmatota Poseidoniia_A Poseidoniales Poseidonia NA NA Poseidoniales Poseidonia
Archaea Thermoplasmatota NA NA NA NA NA Archaea Thermoplasmatota

我认为 paste() 可能适用于这种情况,但不确定如何实现它,因此我可以获得上述预期的输出数据帧。

【问题讨论】:

    标签: r dataframe


    【解决方案1】:

    我们可以使用base R——遍历行,用na.omit删除NA,用n = 2paste得到最后两个元素tail

    demo_table$assignation <- apply(demo_table, 1, 
       function(x) paste(tail(na.omit(x), 2), collapse = " "))
    

    -输出

    demo_table$assignation
    [1] "UBA4441 UBA4441 sp"            "Cyanobacteria Cyanobacteriia"  "Phycisphaerae Phycisphaerales" "Poseidoniales Poseidonia"     
    [5] "Archaea Thermoplasmatota"     
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2020-06-17
      • 1970-01-01
      • 1970-01-01
      • 2021-03-02
      • 1970-01-01
      • 2013-09-03
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多