【问题标题】:Apply self-defined function on list of data frames in R在R中的数据框列表上应用自定义函数
【发布时间】:2020-08-05 08:47:41
【问题描述】:

这是我上一个问题的后续: Code in R to conditionally subtract columns in data frames

我现在想将给定的解决方案应用于我之前的问题

cols <- grep('^\\d+$', names(df), value = TRUE)
new_cols <- paste0(cols, '_corrected')
df[new_cols] <- df[cols] - df[paste0('Background_', cols)]
df[c("Wavelength", new_cols)]

到列表中的每个数据框。我导入了一个 excel 文件的所有工作表,以便每个工作表都使用此代码成为列表中的一个数据框(由 Read all worksheets in an Excel workbook into an R list with data.frames 的最佳答案提供):

read_excel_allsheets <- function(filename, tibble = FALSE) {
  sheets <- readxl::excel_sheets(filename)
  x <- lapply(sheets, function(X) readxl::read_excel(filename, sheet = X))
  if(!tibble) x <- lapply(x, as.data.frame)
  names(x) <- sheets
  x
}

mysheets <- read_excel_allsheets(file.choose())

如何将第一个代码框应用到我的数据框列表?

我想从这样的事情中得到:

df_1 <- structure(list(Wavelength = 300:301, Background_1 = c(5L, 3L), 
                     `1` = c(11L, 12L), Background_2 = c(4L, 5L), `2` = c(12L, 10L)), 
                class = "data.frame", row.names = c(NA, -2L))

df_2 <- structure(list(Wavelength = 300:301, Background_1 = c(6L, 4L),
                     `1` = c(10L, 13L), Background_2 = c(5L, 6L), `2` = c(11L, 11L),
                     Background_3 = c(4L, 6L), `3` = c(13L, 13L)),
                class = "data.frame", row.names = c(NA, -2L))

df_list <- list(df_1, df_2)

这样的:

df_1_corrected <- structure(list(Wavelength = 300:301, `1_corrected` = c(6L, 9L),
                                 `2_corrected` = c(8L, 5L)),
                class = "data.frame", row.names = c(NA, -2L))

df_2_corrected <- structure(list(Wavelength =300:301, `1_corrected` = c(4L, 9L),
                                 `2_corrected` = c(6L, 5L),
                                 `3_corrected` = c(9L, 7L)),
                class = "data.frame", row.names = c(NA, -2L))

df_corrected_list <- list(df_1_corrected, df_2_corrected)

实际数据摘录

Wavelength Background 1        1 Background 2       2 Background 3         3
       300     273290.0 337670.0     276740.0  397530     288500.0  367480.0
       301     299126.7 375143.3     299273.3  432250     310313.3  394796.7

我已经阅读了 lapply 函数将用于此,但我以前从未使用过它,因为我是 R 的初学者。 非常感谢您的帮助!

【问题讨论】:

    标签: r excel list dataframe lapply


    【解决方案1】:

    您可以将代码放在一个函数中,并使用lapply 将其应用于列表中的每个数据框:

    subtract_values <- function(df) {
      cols <- grep('^\\d+$', names(df), value = TRUE)
      new_cols <- paste0(cols, '_corrected')
      df[new_cols] <- df[cols] - df[paste0('Background ', cols)]
      df[c("Wavelength", new_cols)]
    }
    
    lapply(df_list, subtract_values)
    
    #[[1]]
    #  Wavelength 1_corrected 2_corrected
    #1        300           6           8
    #2        301           9           5
    
    #[[2]]
    #  Wavelength 1_corrected 2_corrected 3_corrected
    #1        300           4           6           9
    #2        301           9           5           7
    

    【讨论】:

    • 如果我得到错误我在哪里搞砸了:[.data.frame`(df, paste0("Background", cols)) : undefined columns selected
    • 我想你上次也有同样的错误。它对您共享的测试数据有用吗?您的列名是否与实际数据中显示的相同?
    • 这不是同一个错误,尽管我已经弄清楚,在实际数据中,列名不是“Background_1”而是“Background 1”,带有一个空格。我如何表示它而不是'Background_'?只留下一个空格是不行的
    • 如果你真的有你所说的列名,即“背景 1”,它应该。我也更新了答案。此外,出于完全相同的原因,最好不要在列名中包含空格,因为很难选择它们。
    • 我讨厌听起来没有效率,但它不起作用,我只能通过用 _ 替换数据中的空格并使用您的原始代码来让它工作
    猜你喜欢
    • 2021-12-23
    • 1970-01-01
    • 2018-03-13
    • 1970-01-01
    • 2019-11-09
    • 1970-01-01
    • 2020-11-27
    • 2020-10-27
    • 2021-10-09
    相关资源
    最近更新 更多