【发布时间】:2019-12-28 23:54:05
【问题描述】:
这是一篇重新编辑的帖子 - 原帖不够清晰,但我希望这篇文章得到足够的改进
我有大约 400 个 .csv 文件,所有这些文件的列数都相同(总共 7 个)。每天生成一个文件(因此它们是单独的文件,我希望暂时保持这种方式)。由于获取数据时出现问题,这些文件(大约 30 个左右的连续文件)在其中一个列中缺少数据:Programme_Duration。这些数据很有可能存在于一个或多个其他“完整”文件中(我不会详细说明如何/为什么,但数据中有很多重复)。
以下是相关文件的示例(我不太确定如何共享数据,因为一些观察结果是非常长的字符串。希望这些图片就足够了):
An example of an "incomplete" file, with observations missing for Programme_Duration.
在介绍我正在研究的方法之前,值得指出的是,Programme_Synopsis_url 列将匹配完整文件和不完整文件。因此,这可能是解决这个问题的关键。 IE。创建一个脚本:
- 创建一个包含所有 370 个“完整”csv 文件的数据框(我们称之为 df_complete)。
- 读入第一个“不完整”文件(我们称之为
incomplete_file)。 - 为
Programme_Synopsis_url列标识df_complete和incomplete_file之间的任何匹配行 - 如果匹配,则将
df_complete中Programme_Duration相关行的内容复制到incomplete_file的Programme_Duration对应行中。 - 写出来。
- 重复,即遍历所有 30 个“不完整”文件。
其中一些我可以做到(步骤 1、2、5 和 6!),但重要的中间部分让我很难过。希望这次帖子更清楚。对此的任何帮助将不胜感激!
更新:已解决
如果有人看到这篇文章并处于类似情况,我想我会用我的解决方案更新它。完全披露,这不是一个整洁的解决方案。我几乎没有编程经验,我在下面所做的一切都是通过反复试验和来自不同来源的在线拼凑而成的。我特别感谢@kstew,他的回答帮助我解决了这个问题。
在我分享对我有用的代码之前,还有一个免责声明:这是一个非常独特且不寻常的问题。我认为我在原来的帖子中解释得不是特别好(主要是因为我对这个领域的无知)。事实上,我在执行此操作时还面临其他几个重要因素/挑战 - 例如,我必须保留“不完整”文件中行的原始顺序(我通过简单地添加一个名为 index 的新列解决了这个问题,以便恢复原来的行顺序)。
同样,对于更有经验的程序员来说,这段代码可能看起来一团糟,但它对我有用。话虽如此,请随时编辑/整理!
这是代码,每个步骤都有解释:
### First, create data.frame from "complete" csv files ###
folder_complete <-"insert path here"
df_list_complete <- list.files(path=folder_complete, pattern="*.csv", full.names = TRUE)
df_complete = ldply(df_list_complete, read_csv)
### Then, read in and edit "incomplete" files one at a time using for loop ###
### Note "incomplete" files are in a different director - this was set during the session ###
filenames <- dir(pattern = "*.csv")
for (i in 1:length(filenames)) {
tmp <- read.csv(filenames[i], stringsAsFactors = FALSE)
### Merge / Identify matches between "complete" data.frame and "incomplete"
file ###
### using "Programme Synopsis" as the unique column ###
tmp_new <- merge(tmp, df_complete, by = "Programme_Synopsis")
### Delete any rows with NAs in specific columns - ###
### I did this because the previous step matched empty rows for these columns, and I didn't want these ###
tmp_new <- distinct(tmp_new,Programme_Synopsis_url.x, .keep_all = TRUE)
tmp_new <- distinct(tmp_new,Programme_Duration.y, .keep_all = TRUE)
### Delete Duplicate columns - merging created several duplicate columns (.y, .x) ###
### I only wanted to add the matching "Programme Duration" column from the "complete" data.frame to the "incomplete" file ###
### but wasn't sure how to do this. ###
### Instead, I had to retrospectively remove the duplicate columns ###
tmp_new <- tmp_new[ -c(2:7) ]
### Rename columns ###
tmp_new2 <- rename(tmp_new, c("Programme_Synopsis_url.y" =
"Programme_Synopsis_url",
"Programme_Duration.y" = "Programme_Duration",
"Programme_Category.y" = "Programme_Category",
"Programme_Availability.y" = "Programme_Availability",
"Programme_Genre.y" = "Programme_Genre",
"Programme_Title.y" = "Programme_Title"))
### Merge (again!) using plyr Join function ###
df <- join(tmp_new2, tmp, by = "Programme_Synopsis_url", type = "full")
### Delete any without an index ###
### (i.e. those that don't belong in this dataframe) ###
df <- df[!is.na(df$index), ]
### Re-order by original index ###
df <- df[order(df$index), ]
### Remove duplicated index columns ###
df$index.x <- NULL
df$index.y <- NULL
### Write out the new file ###
write.csv(df, filenames[[i]], row.names = FALSE)
无论如何,希望此更新对其他人有所帮助。同样,如果您看到这个并能想到一个更优雅的解决方案,我很乐意听到它。
【问题讨论】:
-
您最好将每个数据框加载到列表元素中,然后对所有数据框进行连接。 This post 解释了如何合并数据框列表。
标签: r csv merge missing-data