【发布时间】:2016-07-27 10:54:30
【问题描述】:
我想通过读取当月每一天的 csv 文件来构建一个数据框。我的每日 csv 文件包含相同行数的字符、双精度和整数列。我知道任何给定月份的最大行数,并且每个 csv 文件的列数保持不变。我使用 fileListing 循环浏览一个月中的每一天,其中包含 csv 文件名列表(比如一月):
output <- matrix(ncol=18, nrow=2976)
for ( i in 1 : length( fileListing ) ){
df = read.csv( fileListing[ i ], header = FALSE, sep = ',', stringsAsFactors = FALSE, row.names = NULL )
# each df is a data frame with 96 rows and 18 columns
# now insert the data from the ith date for all its rows, appending as you go
for ( j in 1 : 18 ){
output[ , j ] = df[[ j ]]
}
}
很抱歉在我发现部分问题时修改了我的问题(duh),但我应该使用 rbind 逐步在数据框底部插入数据,还是这么慢?
谢谢。
BSL
【问题讨论】:
-
您最好将它们全部读入一个列表,然后使用
do.call(rbind.data.frame, data)一次将它们组合起来。
标签: r csv matrix dataframe pre-allocation