【问题标题】:User defined function for analyzing multiple csvs, create variables, and count NAs by group in R用于分析多个 csv、创建变量和在 R 中按组计算 NA 的用户定义函数
【发布时间】:2020-07-13 05:23:08
【问题描述】:

我有大约 100 个具有不同变量(以及不同数量的变量)的数据集,但每个数据集都有一个家庭 ID (hh_ID) 作为标识符。变量代表调查问题。每个 csv 代表不同类型的调查。我想编写一个自定义函数来计算一个家庭被问到一个问题的次数以及他们跳过一个问题的次数(NA)。我遇到的问题是重命名变量并跨 csvs 计数。

假设两个数据框如下所示:

hh_ID <- c(1,1,2,2,2)
question1 <- c(NA,1,0,0,0)
question2 <- c(1,1,NA,0,0)
df1 <- data.frame(hh_ID, question1, question2)

hh_ID <- c(1,1,1,2,2)
question3 <- c(NA,NA,0,0,0)
question4 <- c(1,1,1,NA,NA)
df2 <- data.frame(hh_ID, question3, question4)

## > df1
##   hh_ID question1 question2
## 1     1        NA         1
## 2     1         1         1
## 3     2         0        NA
## 4     2         0         0
## 5     2         0         0
## > df2
##   hh_ID question3 question4
## 1     1        NA         1
## 2     1        NA         1
## 3     1         0         1
## 4     2         0        NA
## 5     2         0        NA

我需要最终的数据框看起来像这样:

question1_count <- c(2,3)
question1_NAs   <- c(1,0)
question2_count <- c(2,3)
question2_NAs   <- c(0,1)
question3_count <- c(3,2)
question3_NAs   <- c(2,0)
question4_count <- c(3,2)
question4_NAs <- c(0,2)
finaldf <- data.frame(unique(hh_ID),question1_count, question1_NAs,question2_count,question2_NAs,question3_count,question3_NAs, question4_count,question4_NAs) 

## > finaldf
##   unique.hh_ID. question1_count question1_NAs question2_count question2_NAs question3_count question3_NAs question4_count question4_NAs
## 1             1               2             1               2             0               3             2               3             0
## 2             2               3             0               3             1               2             0               2             2

这是我目前所拥有的:

# read in each dta file
filenames <- list.files(path=mydirectory, pattern=".*dta")
for (i in 1:length(filenames)){
assign(filenames[i], read_dta(paste("", filenames[i], sep=''))
)}

variable_NA_count <- function(dataset, col_name){
temp <- dataset %>% group_by(hh_ID) %>% summarise(question_count = n()) 
temp1 <- aggregate(col_name ~ hh_ID, data=dataset, function(x) {sum(is.na(x))}, na.action = NULL)
final <- merge(temp, temp1, by = "hh_ID")
return(final)}

frequency <- function(dataset, col_name){
temp <- variable_NA_count(dataset, col_name)
temp <- temp %>% select(question1_count = question_count,
                        question1_NAs = col_name)}

问题是我希望每个变量名都以“_count”和“_NAs”结尾,而不是明确写“question1_count = question_count”。我在 csvs 中有数百个变量,所以我需要一个函数来读取每个 csv,读取每个列名,计算一个家庭被问到问题的次数,以及他们没有回答的次数。我尝试了各种方法,例如粘贴功能,但一直碰壁。

谢谢!

【问题讨论】:

  • 欢迎来到stackoverflow。生成finaldf似乎有错字。我提供了更正,但请检查它是否符合您的预期。

标签: r function functional-programming


【解决方案1】:

我建议一个快速的解决方案,尽管它与您期望的格式不完全一致。

     list..res <- lapply(list(df1,df2), 
                function(x) setDT(x)[,lapply(.SD,function(x) {  
         list(.N,sum(is.na(x)))}),by=hh_ID][,`:=`(index=1:.N,type=c("count", 
                                                              "no..na")),hh_ID])

对于每个data.frame,我将其转换为data.table(library(data.table)),然后对于每个问题,我计算问题的数量,计算NA 的数量,并计算NA 的数量。最后我添加了一个列type和index

 ## + + > list..res
## [[1]]
##    hh_ID question1 question2 index   type
## 1:     1         2         2     1  count
## 2:     1         1         0     2 no..na
## 3:     2         3         3     1  count
## 4:     2         0         1     2 no..na

## [[2]]
##    hh_ID question3 question4 index   type
## 1:     1         3         3     1  count
## 2:     1         2         0     2 no..na
## 3:     2         2         2     1  count
## 4:     2         0         2     2 no..na

然后我们可以通过合并来减少这个列表。

Reduce(function(x,y) merge(x,y,by=c("hh_ID","type","index")), list..res)

##    hh_ID   type index question1 question2 question3 question4
## 1:     1  count     1         2         2         3         3
## 2:     1 no..na     2         1         0         2         0
## 3:     2  count     1         3         3         2         2
## 4:     2 no..na     2         0         1         0         2

最后,您可以放置​​ data.frames 列表,而不是 list(df1,df2)。

filenames <- list.files(path=mydirectory, pattern=".*dta")
df..list <- lapply(filenames, read_dta)

【讨论】:

    【解决方案2】:

    你可以充分利用dplyr的summarize_all功能:

    它将用一个或多个给定函数汇总df 中的所有列,创建智能列名称(从原始列名称开始并添加函数名称)。

    library(dplyr)
    
    df1 %>%
      group_by(hh_ID) %>% 
      summarize_all(.funs = list(count = ~n(), NAs = ~sum(is.na(.))))
    #> # A tibble: 2 x 5
    #>   hh_ID question1_count question2_count question1_NAs question2_NAs
    #>   <dbl>           <int>           <int>         <int>         <int>
    #> 1     1               2               2             1             0
    #> 2     2               3               3             0             1
    

    由reprex package (v0.3.0) 于 2020-04-01 创建

    我们可以使用purrr 的map 函数将相同的操作应用于数据帧列表:

    library(dplyr)
    library(purrr)
    
    list(df1, df2) %>% 
      map(~{
        .x %>%
          group_by(hh_ID) %>% 
          summarize_all(.funs = list(count = ~n(), NAs = ~sum(is.na(.))))
      }) %>% 
      reduce(full_join)
    #> Joining, by = "hh_ID"
    #> # A tibble: 2 x 9
    #>   hh_ID question1_count question2_count question1_NAs question2_NAs
    #>   <dbl>           <int>           <int>         <int>         <int>
    #> 1     1               2               2             1             0
    #> 2     2               3               3             0             1
    #> # … with 4 more variables: question3_count <int>, question4_count <int>,
    #> #   question3_NAs <int>, question4_NAs <int>
    

    由reprex package (v0.3.0) 于 2020-04-01 创建

    map 返回一个数据框列表,但我们想使用full_join(或您认为合适的任何其他*_join)加入它们

    最后我们可以将它们粘合在一起读取文件:list.files(path=mydirectory, pattern=".*dta") 返回一个字符向量,我们可以将map 应用于它。

    对于每个文件,阅读、总结和加入:

    library(dplyr)
    library(purrr)
    library(haven)
    
    list.files(path=mydirectory, pattern=".*dta") %>% 
      map(~{
        read_dta(.x) %>%
          group_by(hh_ID) %>% 
          summarize_all(.funs = list(count = ~n(), NAs = ~sum(is.na(.))))
      }) %>% 
      reduce(full_join)
    

    由reprex package (v0.3.0) 于 2020-04-01 创建

    (输出不显示,因为我没有任何包含 *.dta 文件的目录)

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-11-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-01-15
      相关资源
      最近更新 更多