【问题标题】:Function to find the (5) highest value of a column and merge together with lowest value according different value of other columns查找一列的(5)最大值并根据其他列的不同值与最小值合并在一起的函数
【发布时间】:2021-08-09 03:36:37
【问题描述】:

我想使用一个函数来更轻松地完成以下任务:通过根据不同的列值分组,在列中找到某个值的 5 个最高值,并将第一列的较低值保留在另一个名称下,并且合并在一起。请阅读下面的脚本以获得初步理解。

我有一个有 5 列的表格: 分类群、年份、站、物种和生物量。 我有 4 种分类群(硅藻、甲藻、鞭毛虫和纤毛虫),两年:2018 年和 2019 年,每年都有不同的站点(2018 年:P1、P2、P3、P4、P5、PICE1、SICE3 和 2019 年:P1、P2 ,P3,P4,P5,P6,P7,Sice4)、不同物种(分类群之间不同)及其各自的生物量值。

我想创建一个数据框:在其中我得到每个分类群的每个物种、每年、每个站点的 5 个最高生物量值。对于其他物种,我想将它们的名称更改为“Others_#”,其中 # 对应于它们各自的分类群(例如:“others_diatoms)。

为此,我已经做了很多步骤,但我确信还有另一种方法。

例如,对于硅藻,我是如何做到的:

  #subset only year 2018, station P1, with decreasing order of the shallowdepth column
P1_18 <- data %>% 
    subset(Taxa == "Diatoms" & Year == "2018" & Station=="P1") %>% 
    arrange(desc(Biomass))
  
  #Keep top 5 value
  top5_P1_18 <- P1_18 %>%
    head(5) 
  
  #keep the rest
  Otherdiatoms_P1_18 <-
    P1_18[-c(1:5),]  #remove the first 5 highest biomass value for each species‘’’
#For P2 in 2018
P2_18 <- data %>% 
  subset(Year== " 2018" & Station=="P2") %>% 
  arrange(desc(Biomass)) 

top5_P2_18 <- P2_18 %>%
  head(5) 

Otherdiatoms_P2_18 <-
  P2_18[-c(1:5),]

每年每个站点的等等,以及其他 3 个分类群。

最后:

#merge top5 tables
top5_diatoms <- bind_rows(top5_P1_19,top5_P2_19,top5_P3_19,top5_P4_19,top5_P5_19,top5_P6_19,top5_P7_19,top5_Sice4_19,
                     top5_P1_18,top5_P2_18,top5_P3_18,top5_P4_18,top5_P5_18,top5_PICE1_18,top5_SICE3_18)

#merge the other diatoms
Others_diatoms<- bind_rows(Otherdiatoms_P1_19,Otherdiatoms_P2_19,Otherdiatoms_P3_19,Otherdiatoms_P4_19,Otherdiatoms_P5_19,Otherdiatoms_P6_19,Otherdiatoms_P7_19,Otherdiatoms_Sice4_19,
                 Otherdiatoms_P1_18,Otherdiatoms_P2_18,Otherdiatoms_P3_18,Otherdiatoms_P4_18,Otherdiatoms_P5_18,Otherdiatoms_PICE1_18,Otherdiatoms_SICE3_18)
#NameSize: change all name species by "Other diatoms"
Others_diatoms$Species <- "Other diatoms"

#combine both together
diatoms_toplot<- bind_rows(top5_diatoms, Others_diatoms)

and then #combine the 4 tables taxa (

Final_merge<-bind_rows(diatoms_toplot, dinoflagellates_toplot,flagellates_toplot,ciliates_toplot)

我正在考虑为数据框的几列创建一个列表,并在每个列上应用一个函数......但我有点迷茫,所以如果有任何帮助,即使是小步骤,我也会很高兴: )

我的一小部分数据示例: dput(data) structure(list(Taxa = c("Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Diatoms", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Ciliates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Flagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates", "Dinoflagellates" ), Year = c(2018, 2019, 2019, 2018, 2018, 2019, 2019, 2019, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2019, 2018, 2018, 2019, 2018, 2019, 2019, 2018, 2018, 2019, 2019, 2019, 2018, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2018, 2018, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2019, 2019, 2019, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2018, 2018, 2019, 2018, 2018, 2019, 2018, 2019, 2018, 2018, 2018, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2019, 2018, 2019, 2019, 2019, 2019, 2019, 2019, 2019, 2018, 2019, 2018, 2019, 2018, 2019, 2019, 2018, 2019, 2019, 2019, 2018, 2019, 2019, 2019, 2019, 2018, 2018, 2019, 2018), Station = c("P3", "P6", "P7", "PICE1", "SICE3", "P4", "P5", "P6", "P7", "SICE3", "Sice4", "P1", "P1", "P2", "P3", "P5", "P6", "P7", "PICE1", "SICE3", "Sice4", "SICE3", "P4", "P7", "PICE1", "SICE3", "Sice4", "P5", "P7", "PICE1", "Sice4", "P4", "P5", "P7", "P3", "P4", "P6", "P1", "P2", "P3", "P5", "P7", "Sice4", "P1", "Sice4", "P3", "P1", "P2", "P2", "P5", "P7", "PICE1", "Sice4", "P2", "P2", "P3", "P3", "P5", "P6", "P7", "Sice4", "P1", "P1", "P2", "P2", "P3", "P3", "P4", "P4", "P5", "SICE3", "P1", "P2", "PICE1", "SICE3", "Sice4", "P1", "Sice4", "P1", "P1", "P2", "P2", "P3", "P3", "P4", "P4", "P5", "P5", "P6", "P3", "P3", "P4", "P5", "P6", "P7", "Sice4", "P1", "P2", "P2", "P3", "P3", "P4", "P4", "P2", "P3", "P3", "P4", "Sice4", "P2", "P2", "P5", "P6", "P7", "PICE1", "SICE3", "Sice4", "SICE3" ), Species = c("Chaetoceros hibiscus", "Chaetoceros hibiscus", "Chaetoceros hibiscus", "Chaetoceros hibiscus", "Chaetoceros hibiscus", "Chaetoceros australis", "Chaetoceros australis", "Chaetoceros australis", "Chaetoceros australis", "Chaetoceros australis", "Chaetoceros australis", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Cylindrotheca", "Entomoneis", "Eucampia", "Eucampia", "Eucampia", "Eucampia", "Eucampia", "Fragilariopsis", "Fragilariopsis", "Fragilariopsis", "Fragilariopsis", "Fragilariopsis nana", "Fragilariopsis nana", "Fragilariopsis nana", "Bicosta spinifera", "Bicosta spinifera", "Bicosta spinifera", "Pathronus", "Pathronus", "Pathronus", "Pathronus", "Pathronus", "Pathronus", "Plusimus", "Mesodinium", "Scuticociliatia", "Braconus", "Braconus", "Braconus", "Braconus", "Braconus", "Braconus", "Braconus", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Acanthostomella", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Liboa", "Leegaardiella ovalis", "Leegaardiella ovalis", "Leegaardiella ovalis", "Leegaardiella ovalis", "Leegaardiella ovalis", "Leegaardiella sol", "Leegaardiella sol", "Leprotintinnus", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Lohmannio", "Chrysophyceae", "Chrysophyceae", "Chrysophyceae", "Chrysophyceae", "Chrysophyceae", "Chrysophyceae", "Chrysophyceae", "Dinobryonus", "Dinobryonus", "Dinobryonus", "Dinobryonus", "Dinobryonus", "Dinobryonus", "Dinobryonus", "Amphidus", "Amphidus", "Amphidus", "Amphidus", "Amphidus", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidipoma", "Amphidoma acuminata"), Biomass = c(2.5722570760978, 106.489043401096, 11.3660482634744, 23.6180471200604, 4.09736836295174, 2.77106585864742, 12.223960073548, 72.7319818648167, 8.1427752718168, 3.58881675476654, 101.787835667047, 9.47093811201888, 0.14002298927298, 0.0767060109516124, 0.0373241446028316, 0.169776162765161, 1.08282241487347, 0.195956004017996, 0.574227994044988, 1.30410474481671, 0.362778732033109, 5.11147483254336, 7.67526064939454, 11.1631040712476, 31.3527468551935, 58.6443407164168, 90.2661658419651, 0.0446244917247616, 0.840336011374516, 0.652638825938322, 1.60247037093698, 1.89950606550275, 13.9012274189626, 1.03046243784118, 0.966327867145711, 0.316060682764314, 0.165776507451685, 0.58830523425, 28.166516349, 9.36453636, 1.019766267, 0.9057995505, 2.6816984875, 0.144669507, 29.201684359, 46.8677637926364, 13.6932074914416, 24.4990583557351, 1.74088105245273, 3.57267406669825, 0.399665183658365, 18.7494388294602, 1.65382368830652, 1.68834080413432, 1.67395923, 1.7191065981177, 0.25746012, 3.42252693, 10.46782755, 12.63325329, 27.28561977, 44.3880970044607, 9.62682960654306, 54.4142437051378, 26.354896524942, 41.6469905251234, 31.7857807556355, 18.7536292316292, 10.7666063235056, 48.0379878500564, 9.08744172316435, 22.7274163346186, 3.65192443308688, 2.04894292586637, 2.25383716895731, 12.0194344079741, 447.459734842254, 46.5926400106843, 6.37144904526355, 4.60879900929901, 5.52746175351058, 0.763267538, 4.75285289064043, 0.167946768, 1.00104490733456, 0.376861628, 32.4558052933942, 1.661617478, 3.61765976, 37.45879402342, 63.7914816, 12.4429536, 73.16948865, 2.2097162, 8.4491484, 0.87203655, 5.56556268491627, 0.042503908636136, 12.8078492514965, 0.370958637607761, 0.970623318853854, 0.251205938893951, 2.95334285993973, 1.1088252, 1.16333569359, 0.697977, 0.2501982, 8.119062, 1.0277386590735, 0.26885925, 0.3437478, 11.4169177, 0.0743094, 0.5457011783256, 0.36260916888235, 1.62785205, 0.0876508754603 )), row.names = c(NA, -117L), class = c("tbl_df", "tbl", "data.frame" ))

【问题讨论】:

  • 是否有两个数据集,即datadepthint_biomass_5sp_diatoms
  • 只有1个数据集,名称为“数据”,错误会添加示例...
  • 好的,请您更新您的帖子。由于有几个子集,有点混乱

标签: r list function loops


【解决方案1】:

我们可以在分组后使用slice_head

library(dplyr)
data %>% 
   arrange(Taxa, Year, Station, desc(Biomass)) %>%
   group_by(Taxa, Year, Station) %>%
   slice_head(n = 5)

slice_max

data %>%
    group_by(Taxa, Year, Station) %>%
    slice_max(n = 5, order_by = Biomass)

如果我们需要将“Taxa”的名称更改为前缀为“Others”的那些不是前 5 名的名称

library(stringr)
df1 <- data %>%
  arrange(Taxa, Year, Station, desc(Biomass)) %>%
  group_by(Taxa, Year, Station) %>%
  slice_head(n = 5) 
nm1 <- unique(df1$Taxa)
data <- data %>%
    mutate(Taxa = case_when(! Taxa %in% nm1 ~ str_c('Other_', Taxa), TRUE ~ Taxa))

然后,我们 filter 已经创建了“其他”以及“df1”的那些

data %>%
     filter(str_detect(Taxa, '^Other_')) %>%
     bind_rows(df1 %>% ungroup)

【讨论】:

  • 好的,非常感谢。是否可以创建一个 df 但不是保留 5 first ,而是先删除五个?我试过:数据 %>% 安排(Taxa, Year, Station, desc(Biomass)) %>% group_by(Taxa, Year, Station) %>% slice((6:length(data1))) 或者其他的想法是从原始表中减去保存的行(从 head = 5 函数)(因为它们是重复的):df1 %arrange(Taxa, Year, Station, desc(Biomass)) %>% group_by (Taxa, Year, Station) %>% slice_head(n = 5) and data - df1
  • 问题已解决:#keep 5 个最高的生物量 data_head5 % 排列(Taxa, Year, Station, desc(Biomass)) %>% group_by(Taxa, Year, Station) %> % slice_head(n = 5) #得到没有data1行的df data_without5head
猜你喜欢
  • 2021-07-25
  • 2021-08-21
  • 2021-05-17
  • 1970-01-01
  • 2015-11-28
  • 1970-01-01
  • 2015-04-22
  • 2021-05-25
  • 1970-01-01
相关资源
最近更新 更多