【问题标题】:(R) Comparing multiple csv files and adding in missing values/observations(R) 比较多个 csv 文件并添加缺失值/观察值
【发布时间】:2019-12-28 23:54:05
【问题描述】:

这是一篇重新编辑的帖子 - 原帖不够清晰,但我希望这篇文章得到足够的改进

我有大约 400 个 .csv 文件,所有这些文件的列数都相同(总共 7 个)。每天生成一个文件(因此它们是单独的文件,我希望暂时保持这种方式)。由于获取数据时出现问题,这些文件(大约 30 个左右的连续文件)在其中一个列中缺少数据:Programme_Duration。这些数据很有可能存在于一个或多个其他“完整”文件中(我不会详细说明如何/为什么,但数据中有很多重复)。

以下是相关文件的示例(我不太确定如何共享数据,因为一些观察结果是非常长的字符串。希望这些图片就足够了):

An example of a "complete" file, with observations for Programme_Duration (note: not every row has an observation).

An example of an "incomplete" file, with observations missing for Programme_Duration.

在介绍我正在研究的方法之前,值得指出的是,Programme_Synopsis_url 列将匹配完整文件和不完整文件。因此,这可能是解决这个问题的关键。 IE。创建一个脚本:

  1. 创建一个包含所有 370 个“完整”csv 文件的数据框(我们称之为 df_complete)。
  2. 读入第一个“不完整”文件(我们称之为incomplete_file)。
  3. 为Programme_Synopsis_url 列标识df_complete 和incomplete_file 之间的任何匹配行
  4. 如果匹配,则将df_complete中Programme_Duration相关行的内容复制到incomplete_file的Programme_Duration对应行中。
  5. 写出来。
  6. 重复,即遍历所有 30 个“不完整”文件。

其中一些我可以做到(步骤 1、2、5 和 6!),但重要的中间部分让我很难过。希望这次帖子更清楚。对此的任何帮助将不胜感激!

更新:已解决

如果有人看到这篇文章并处于类似情况,我想我会用我的解决方案更新它。完全披露,这不是一个整洁的解决方案。我几乎没有编程经验,我在下面所做的一切都是通过反复试验和来自不同来源的在线拼凑而成的。我特别感谢@kstew,他的回答帮助我解决了这个问题。

在我分享对我有用的代码之前,还有一个免责声明:这是一个非常独特且不寻常的问题。我认为我在原来的帖子中解释得不是特别好(主要是因为我对这个领域的无知)。事实上,我在执行此操作时还面临其他几个重要因素/挑战 - 例如,我必须保留“不完整”文件中行的原始顺序(我通过简单地添加一个名为 index 的新列解决了这个问题,以便恢复原来的行顺序)。

同样,对于更有经验的程序员来说,这段代码可能看起来一团糟,但它对我有用。话虽如此,请随时编辑/整理!

这是代码,每个步骤都有解释:


### First, create data.frame from "complete" csv files ###
folder_complete <-"insert path here"
df_list_complete <- list.files(path=folder_complete, pattern="*.csv", full.names = TRUE)
df_complete = ldply(df_list_complete, read_csv)

### Then, read in and edit "incomplete" files one at a time using for loop ### 
### Note "incomplete" files are in a different director - this was set during the session ###
filenames <- dir(pattern = "*.csv")
for (i in 1:length(filenames)) {
    tmp <- read.csv(filenames[i], stringsAsFactors = FALSE)
    ### Merge / Identify matches between "complete" data.frame and "incomplete" 
    file ### 
    ### using "Programme Synopsis" as the unique column ###
    tmp_new <- merge(tmp, df_complete, by = "Programme_Synopsis")
    ### Delete any rows with NAs in specific columns - ###
    ### I did this because the previous step matched empty rows for these columns, and I didn't want these ###
    tmp_new <- distinct(tmp_new,Programme_Synopsis_url.x, .keep_all = TRUE)
    tmp_new <- distinct(tmp_new,Programme_Duration.y, .keep_all = TRUE)
    ### Delete Duplicate columns - merging created several duplicate columns (.y, .x) ###
    ### I only wanted to add the matching "Programme Duration" column from the "complete" data.frame to the "incomplete" file ###
    ### but wasn't sure how to do this. ###
    ### Instead, I had to retrospectively remove the duplicate columns ###
    tmp_new <- tmp_new[ -c(2:7) ]
    ### Rename columns ###
    tmp_new2 <- rename(tmp_new, c("Programme_Synopsis_url.y" = 
    "Programme_Synopsis_url", 
    "Programme_Duration.y" = "Programme_Duration",
    "Programme_Category.y" = "Programme_Category", 
    "Programme_Availability.y" = "Programme_Availability", 
    "Programme_Genre.y" = "Programme_Genre", 
    "Programme_Title.y" = "Programme_Title"))
    ### Merge (again!) using plyr Join function ###
    df <- join(tmp_new2, tmp, by = "Programme_Synopsis_url", type = "full")
    ### Delete any without an index ###
    ### (i.e. those that don't belong in this dataframe) ###
    df <- df[!is.na(df$index), ]
    ### Re-order by original index ###
    df <- df[order(df$index), ]
    ### Remove duplicated index columns ###
    df$index.x <- NULL
    df$index.y <- NULL
    ### Write out the new file ###
    write.csv(df, filenames[[i]], row.names = FALSE)

无论如何,希望此更新对其他人有所帮助。同样,如果您看到这个并能想到一个更优雅的解决方案,我很乐意听到它。

【问题讨论】:

  • 您最好将每个数据框加载到列表元素中,然后对所有数据框进行连接。 This post 解释了如何合并数据框列表。

标签: r csv merge missing-data


【解决方案1】:

我会采取与 Gautam 不同的方法。由于您的 CSV 文件具有相同数量的列,因此您可以一次读取所有文件,使用它们的文件名(假设这是按日期唯一的)作为排序 ID,并将 unnest 它们放入一个大数据框中。然后,您可以通过所需的列(在这种情况下为持续时间)过滤掉不完整的案例,并根据完整案例中存在的数据和匹配列(即 URL)替换缺失值。

# list CSV files in your directory
l <- list.files('./','.*csv',full.names = F)

# read in files, all cols as character, and map to one df
df <- data.frame(filename=l) %>% 
  mutate(cont = map(filename,
                    ~ read_csv(file.path(.),col_types = cols(.default = 'c')))) %>% 
  unnest(.)

# find cases with missing duration
incomp <- df %>% mutate(duration=as.numeric(duration)) %>% 
  filter(is.na(duration)) %>% rename(duration.old=duration)

# find cases with present duration
comp <- df %>% mutate(duration=as.numeric(duration)) %>% 
  filter(!is.na(duration))

# replace missing duration based on matching complete cases
comp %>% distinct(filename,day,url,duration) %>% left_join(incomp,.)

  filename day url duration.old duration
1 day2.csv   2   i           NA       22
2 day6.csv   6   j           NA       26
3 day6.csv   6   j           NA       12
4 day6.csv   6   j           NA       16

然后,您可以使用下面的d_ply 方法一个一个地写出新“完成”的文件。

# data used
set.seed(123)
df <- data.frame(day=sample(1:10,100,T),
                 url=sample(letters[1:10],100,T),
                 duration=sample(c(10:30,NA),100,T))
d_ply(df,.(day),function(x) write_csv(x,paste0('day',x %>% distinct(day),'.csv')))

【讨论】:

  • 谢谢,我认为这种方法可能会奏效(@Gautams 也一样——正如他们所说,给猫剥皮的方法不止一种!)不过我可能需要稍微修改一下——我的帖子可能会对Programme_Duration 列中的缺失值更加清楚。此列中的值并非在每一行中都存在,实际上它们在“完整”文件中偶尔出现。我感觉您的方法(搜索任何缺失值的实例)因此可能需要以某种方式进行修改。希望这有意义吗?不管怎样,我可能要到下周才有机会尝试这个……
  • 再想一想,我想我刚刚理解了你的方法。你建议我应该搜索所有缺失的值,然后搜索这些值是否存在于其他地方,如果存在,将它们插入到缺失的行中?顺便说一句,如果这使整个过程更容易,将完整和不完整的文件保存到不同的目录是很容易的?
  • 嗨@Japes,我认为你的第一个问题是肯定的。据我了解您的 OP 中的问题,您只需将某些列中的缺失值替换为其他地方已经存在的值,匹配某些所需的标准(即,如果 URL 相同,则将 NA 替换为该 URL 的程序持续时间) .如果是这种情况,您不一定需要先将文件分类为完整/不完整。替换缺失值后,您可以继续工作流程的其余部分。
  • 在我看来,首先将文件视为完整/不完整并标记它们是不必要的步骤。除非我误解了您的预期输出(例如,在您替换缺失值后覆盖不完整的文件)?
  • 是的,你已经明确了我想要实现的目标。我没有编程背景(我是一名电视学者!)所以我会在这里向你更好的判断低头。我认为将完整/不完整的文件分隔到不同的目录会有所帮助,但我现在可以看到它只会添加一个额外的步骤。感谢您的澄清。一旦有机会,我会尝试一下,并会告诉你我的进展情况
【解决方案2】:

写作为答案,因为评论太长了。

当我有大型数据集或必须读取大量文件时,我更喜欢使用data.table,它速度快且内存效率高。下面的代码将所有“完整”文件读入一个列表,然后将该列表合并为一个大data.table:

代码

library(data.table)

# assuming you have all csv files in the same location
basedir <- choose.dir()
fnames <- dir(path = basedir, pattern = '.*csv', all.files = T, full.names = T, recursive = F)

# names for the columns you want to read, using alphabets as an example 
column_names <- LETTERS[1:7]

big_list <- lappply(fnames, function(fname){
  dat <- fread(file = fname, select = 1:7, col.names = column_names)

  # test for empty column, say, column B 
  if( dat[!is.na(B), .N] < nrow(dat)){
    dat$type <- 'imcomplete'
  }else{
    dat$type <- 'complete'
  }

})

# combine them all into one list
big_data <- rbindlist(l = big_list, use.names = T, fill = T)

# set column B as the key
setkey(big_data, 'B')

complete <- big_data[type == 'complete']
incomplete <- big_data[type == 'incomplete']

或者,您可以根本不读取“不完整”的文件。

对于填写缺失的数据,您可以使用多种不同的方式,我不确定您要使用的逻辑。例如,您可以 merge 仅使用不同的 complete 数据集的不完整部分。

此处描述了基于key 的data.table 方法:Fill missing values from another dataframe with the same columns(单行)。我不确定你想使用什么逻辑,所以我在这里编造一些:

# sample for complete
complete <- as.data.table(mtcars)

# sample for incomplete
incomplete <- as.data.table(mtcars[1:20, ])

# set some values to NA - examples of missing data
incomplete[runif(5, 1, .N), mpg := NA]

验证:

> incomplete[is.na(mpg), .N]
[1] 5

基于wt和qsec合并:

# I used the on argument on two variables 
# because both wt and qsec are not unique for each observation
setDT(incomplete)[complete, mpg := i.mpg, on = .(wt, qsec)]

结果:

> incomplete[is.na(mpg), .N]
[1] 0

# Verifying that we have the right values filled in
> identical(complete[1:20, ], incomplete)
[1] TRUE

要写出结果,您可以使用fwrite(complete, 'complete.csv')。如果需要,您可以省略 type 列。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2017-08-28
    • 2021-05-08
    • 2012-01-09
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-04-04
    相关资源
    最近更新 更多