【问题标题】:Progressive appending of data from read.csv从 read.csv 逐步追加数据
【发布时间】:2016-07-27 10:54:30
【问题描述】:

我想通过读取当月每一天的 csv 文件来构建一个数据框。我的每日 csv 文件包含相同行数的字符、双精度和整数列。我知道任何给定月份的最大行数,并且每个 csv 文件的列数保持不变。我使用 fileListing 循环浏览一个月中的每一天,其中包含 csv 文件名列表(比如一月):

output <- matrix(ncol=18, nrow=2976)
for ( i in 1 : length( fileListing ) ){
    df = read.csv( fileListing[ i ], header = FALSE, sep = ',', stringsAsFactors = FALSE, row.names = NULL )
    # each df is a data frame with 96 rows and 18 columns

    # now insert the data from the ith date for all its rows, appending as you go
        for ( j in 1 : 18 ){        
            output[ , j ]   = df[[ j ]]
        }
}

很抱歉在我发现部分问题时修改了我的问题(duh),但我应该使用 rbind 逐步在数据框底部插入数据,还是这么慢?

谢谢。

BSL

【问题讨论】:

  • 您最好将它们全部读入一个列表,然后使用do.call(rbind.data.frame, data) 一次将它们组合起来。

标签: r csv matrix dataframe pre-allocation


【解决方案1】:

如果数据相对于您的可用内存而言相当小,只需将数据读入即可,不必担心。在您读入所有数据并进行一些清理后,使用 save() 保存文件并使用 load() 让您的分析脚本读取该文件。将读取/清理脚本与分析剪辑分开是减少此问题的好方法。

加快读取 read.csv 的一个功能是使用 nrow 和 colClass 参数。既然你说你知道每个文件中的行数,告诉 R 这将有助于加快阅读速度。您可以使用

提取列类
colClasses <- sapply(read.csv(file, nrow=100), class)

然后将结果提供给 colClass 参数。

如果数据接近太大,您可以考虑处理单个文件并保存中间版本。网站上有许多关于管理内存的相关讨论,涵盖了这个主题。

关于内存使用技巧: Tricks to manage the available memory in an R session

关于使用垃圾收集器功能: Forcing garbage collection to run in R with the gc() command

【讨论】:

  • 我想到了那一步,但还是想写每日文件的每月集合,这样在每月数据框的第1天数据的底部追加第2天。谢谢。
  • 对 colClass 和 nrow 参数进行一些编辑。这些将有助于读取时间和内存使用。在中等大小的数据集上使用 rbind 会很快。
【解决方案2】:

您可以使用lapply 将它们读入一个列表,然后一次将它们组合起来:

data <- lapply(fileListing, read.csv, header = FALSE, stringsAsFactors = FALSE, row.names = NULL)
df <- do.call(rbind.data.frame, data)

【讨论】:

    【解决方案3】:

    首先定义一个主数据框来保存所有数据。然后在读取每个文件时,将数据附加到主服务器上。

    masterdf<-data.frame()
    for ( i in 1 : length( fileListing ) ){
      df = read.csv( fileListing[ i ], header = FALSE, sep = ',', stringsAsFactors = FALSE, row.names = NULL )
      # each df is a data frame with 96 rows and 18 columns
      masterdf<-rbind(masterdf, df)
    }
    

    在循环结束时,masterdf 将包含所有数据。此代码代码可以改进,但对于数据集的大小,这应该足够快。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2015-04-16
      • 1970-01-01
      • 2021-12-21
      • 1970-01-01
      • 2014-04-11
      • 2014-11-21
      • 1970-01-01
      相关资源
      最近更新 更多