【问题标题】:How to put a loop on an sapply function to create a subset for multiple participants?如何在 sapply 函数上放置一个循环来为多个参与者创建一个子集?
【发布时间】:2018-06-09 09:36:18
【问题描述】:

对于以下问题,我将不胜感激:

我有多个巨大的日志文件(每个 >1.000.000 个条目),其中包含一些我特别感兴趣的行(行)。所以我想制作一个仅包含这些行的子集,但我想将结果写入包含多个日志文件/参与者信息的矩阵中。因此,我创建了一小行代码来 1. 创建子集并 2. 在循环中运行它,不仅针对其中一个日志文件,而且针对所有日志文件。

  Result <- subset(df, df$columnOfInterest== "interestingCondition1" | df$columnOfInterest== "interestingCondition2" | df$columnOfInterest== "interestingCondition3")$columnOfInterest
  View(Result)

1
interestingCondition1
2
interestingCondition1
3
interestingCondition2
4
interestingCondition1
5
interestingCondition1
6
interestingCondition3
7
interestingCondition2
8
interestingCondition1
9
interestingCondition1
10
interestingCondition1

嵌入到循环中:

WrongResult <- matrix(data=NA,nrow=TrialNumber, ncol=length(ListOfFiles))
vpncount <- 1
for (v in ListOfFiles){

  df<- read.delim(v, header = TRUE, sep='\t')
  WrongResult[,vpncount] <- subset(df, df$columnOfInterest== "interestingCondition1" | df$columnOfInterest== "interestingCondition2" | df$columnOfInterest== "interestingCondition3")$columnOfInterest

vpncount <- vpncount+1

}

在一个日志文件上运行代码时,我得到了我想要的结果,但是当通过循环运行它时,它会创建一个具有适当大小的矩阵,但只是填充了“随机”数字而不是我细分的条件.

有谁知道为什么会发生这种情况以及如何解决它?非常感谢任何帮助!

编辑:

我尝试创建一个示例数据框。第一行代码(包括变量 Results)就像我想要的那样工作。它在我的 columnOfInterest 的行上过滤我的数据框并将它们放入一个新矩阵中。但是,如果我尝试在一个循环中运行它不止一个数据帧,我就会一直遇到错误:

df <- data.frame(
  X = sample(1:10),
  columnOfInterest= sample(c("interestingCondition1", "interestingCondition2", "interestingCondition3", "NotinterestingCondition1"), 10, replace = TRUE)
)

View(df)

Result <- subset(df, df$columnOfInterest== "interestingCondition1" | df$columnOfInterest== "interestingCondition2" | df$columnOfInterest== "interestingCondition3")$columnOfInterest
View(Result)

WrongResult <- matrix(data=NA,nrow=280, ncol=20)
vpncount <- 1
for (v in 1:20){

  df<- read.delim(v, header = TRUE, sep='\t')
  WrongResult[,vpncount] <- subset(df, df$columnOfInterest== "interestingCondition1" | df$columnOfInterest== "interestingCondition2" | df$columnOfInterest== "interestingCondition3")$columnOfInterest

  vpncount <- vpncount+1

}

View(WrongResult)

【问题讨论】:

  • 我没有看到你的 sapply 函数?
  • 是的,你说得对,我真的很抱歉。我测试了几个解决方案,上面我用子集()替换了 sapply 行。这是另一个想法,对我来说也很有效,但是将其放入循环(因此在多个数据帧上运行)的问题仍然存在。
  • WrongResult2
  • 请更正样本数据中的错别字 -- columnOfInterest, interestingCondition1
  • @MartinMorgan 完成,对不起!

标签: r loops matrix sapply


【解决方案1】:

有人知道为什么会这样吗?

你的循环……不工作。原因有点复杂,但我在 base R 中使用简单的循环(没有 *apply 函数)做了一个工作示例,希望您能继续跟进,并希望它在一定程度上代表您的问题。

在跑步之前先学会走路。在学习如何更简洁地使用apply()、lapply() 等之前学习基本循环。在深入研究非标准评估(data.table、@ 987654324@、purrr等)

首先我们将创建一些数据帧并将它们写入文件

owd <- getwd()
dir.create("sotest")
setwd("sotest")

set.seed(1)

flist <- c("dtf1.txt", "dtf2.txt", "dtf3.txt")

for (i in 1:length(flist)) {
    dtf <- data.frame(
      X=sample(1:10),
      coi=sample(c("ic1", "ic2", "ic3", "nic1"), 10, replace=TRUE)
    )
    write.table(dtf, flist[i], row.names=FALSE, sep="\t")
}

运行此程序后,您应该有一个名为“sotest”的文件夹,其中包含三个制表符分隔的 txt 文件。

然后我们会得到一个可用文件列表,然后循环遍历它。

flist <- list.files(pattern=".txt")
WrongResult <- list()
interesting <- c("ic1", "ic2", "ic3")

for (v in 1:length(flist)) {

    dtf <- read.delim(flist[v], header=TRUE, sep="\t", stringsAsFactors=FALSE)
    WrongResult[[v]] <- dtf[dtf$coi %in% interesting, "coi"]

}

WrongResult

setwd(owd)

我将输出存储为列表而不是矩阵,因为循环的每次迭代中生成的对象的长度不一样。

【讨论】:

  • 非常感谢!您的解决方案对我来说非常有效,我非常感谢您花费了一些额外的时间和精力向初学者解释它,并且采用了一种不涉及很多新软件包的方法。我可以很好地遵循它并将其用于我的问题。谢谢一百万!
【解决方案2】:

我不记得如何使用 data.frame 了,所以我会尝试使用 data.table。你可能需要安装 data.table 包以防你没有它install.packages("data.table")

library(data.table)
dt <- data.table(df)

然后你可以通过以下方式重写你的代码

subset..table <- function(dt){
    dt[columnOfInterest %in% c("interestingCondition1",
                               "interestingCondition2",
                               "interestingCondition3"),columnOfInterest]
}


myfun <- function(x){
### DD
    ## x interp string representing  file name

### Purpose
    ## read and subset

    dt <- fread(x,header=TRUE,sep="\t")
    subset..table(dt)

}

res..list <- lapply(ListOfFiles, myfun)

编辑

例如使用您的示例。

df <- data.frame(
  X = sample(1:10),
  columnOfInterest= sample(c("interestingCondition1",
    "interestingCondition2", "interestingCondition3", 
    "NotinterestingCondition1"), 10, replace = TRUE))


dt <- data.table(df)
subset..table(dt)

会产生

#[1] "interestingCondition2" "interestingCondition3" "interestingCondition1"
#[4] "interestingCondition2" "interestingCondition1" "interestingCondition2"
#[7] "interestingCondition3" "interestingCondition1" "interestingCondition3"

如果您对函数子集..table 感到满意,那么您只需要使用函数 myfun 即可得到您想要的。函数fread 会自动给你一个data.table。

【讨论】:

  • 我只看到 fread() 依赖于 data.table。你不能用 OP 使用的read.delim() 替换它吗?
  • 不幸的是,这不起作用。错误消息告诉我们,未找到 columnOfInterest(对象)。任何想法,为什么会这样?
  • 怎么样?我在那里没有看到任何特定于 data.table 的内容。
  • @AlexB.:如果你给我们reproducible example,那么给你一个可行的解决方案会更容易。
  • @DJJ columnOfInterest 是 code df code 的一列。 @AkselA 我正在研究一个可重现的示例,因为我的数据框太大而无法直接复制。
【解决方案3】:

在 tidyverse 领域中,当您处理单个数据帧时,您希望 filter() 然后 select() 您的原始数据,并为方便起见,使用 mutate() 添加文件名。当有多个可能的值时,一个很好的过滤方法是使用%in%。所以

library(tidyverse)

process_1_df <- function(df, id, condition)
    select(df, columnOfInterest) %>%                 # only interesting column
        filter(columnOfInterest %in% condition) %>%  # specific rows
        mutate(id = id)                              # add identifier

condition <- paste0("interestingCondition", 1:3)
process_1_df(df, "id", condition)

id 是一个标识符——如果 data.frame 来自文件 'foo.txt',则使用 "foo.txt" 作为 id。最初的问题试图将来自多个文件的数据表示为一个矩阵,但假设每个文件都选择了相同数量的有趣行。这里的策略是创建一个数据框,其中包含感兴趣条件来自的文件以及感兴趣条件的值。这个数据框在处理多个文件时很有用...

这适用于样本数据集:

> condition <- paste0("interestingCondition", 1:3)
> process_1_df(df, "id", condition)
       columnOfInterest id
1 interestingCondition2 id
2 interestingCondition2 id
3 interestingCondition3 id
4 interestingCondition1 id
5 interestingCondition3 id
6 interestingCondition1 id

你可以扩展它来处理一个文件

process_1_file <- function(file_name, condition)
    read_csv(file_name) %>%                   # better: input only columnOfInterest
        process_1_df(file_name, condition)

正如@DJJ 所建议的,process_1_file() 的 data.table 实现可能非常紧凑和高效——fread(file_name)[columnOfInterest %in% condition, columnOfInterest]

要处理多个文件,请使用 purr 包

library(purrr)
process_files <- function(file_names, condition)
    map(file_names, process_1_file, condition) %>%
        bind_rows()

dir(pattern="*.csv") %>% process_files(condition)

最终结果是一个数据框,其中有一列感兴趣的条件,另一列指示感兴趣的条件来自哪个日志文件。现在可以根据需要处理/汇总这个“长”格式的数据帧。

【讨论】:

  • 非常感谢您的努力!我尝试使用该代码,但未过滤结果。生成的矩阵由提供的示例数据帧的所有 10 行组成,并且 NotinterestingCondition1 仍然包括在内。我不能真正遵循您的代码,因为提到的包和 R 通常对我来说很新。你现在知道为什么会这样了吗?
  • 您是否确保已纠正错别字?我更新了我的帖子以证明它有效。
  • 复制您更新的代码我仍然收到以下错误消息:'filter_impl(.data, quo) 中的错误:评估错误:缺少参数“条件”,没有默认值。'
  • 您是否定义了条件变量,condition &lt;- paste0("interestingCondition", 1:3) 表示感兴趣的条件?
  • 是的,我做到了,它正常工作:condition [1] "interestingCondition1" "interestingCondition2" "interestingCondition3。该功能仍然无法正常工作:process_1_df(df, condition) Error in filter_impl(.data, quo) : Evaluation error: argument "condition" is missing, with no default.
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-04-16
相关资源
最近更新 更多