【发布时间】:2016-06-11 23:32:28
【问题描述】:
这是我文件夹中的三个示例数据文件(文本文件)。我试图提取每个数据文件对应的日期,并为所有三个文件创建一个数据框。 以下是文件:
文件1:
### DEFAULTS ###
### OPTIONS ###
Title 2 left $ Date and time: 2010-02-03 12:00 UTC
### DATA ###
250 23.54300 0.00000 12.90000
500 17.47400 3.50000 21.70000
750 41.33200 22.10000 30.40000
文件2:
### DEFAULTS ###
### OPTIONS ###
Title 2 left $ Date and time: 2010-02-10 12:00 UTC
### DATA ###
250 36.95300 30.60000 27.10000
500 40.87700 37.80000 27.80000
750 41.46100 38.30000 30.70000
文件3:
### DEFAULTS ###
### OPTIONS ###
Title 2 left $ Date and time: 2010-02-17 12:00 UTC
### DATA ###
250 28.91200 24.90000 5.00000
500 49.82900 32.40000 11.10000
750 50.83600 40.60000 22.20000
下面是我为提取日期编写的代码。
### Start R Code for extracting date
setwd("C:/Users/")
path = "~C:/Users"
file.names<- dir("C:/Users/", pattern =".dat")
file.names
fileNameVector <- NULL
dateNameVector <- NULL
for(i in 1:length(file.names)){
x<-readLines(file.names[i])
x
date.list<-x[3]
date.list
xx<-strsplit(date.list,' ')
xx
date.fin<-unlist(xx)[13]
fileNameVector <- rbind(fileNameVector, file.names[i])
dateNameVector <- rbind(dateNameVector, date.fin)
}
finalDateName <- cbind(fileNameVector, dateNameVector)
finalDateName
用于提取值:
setwd("C:/Users")
path = "~C:/Users"
st021ozone<-""
file.names<- dir("C:/Users", pattern =".dat")
for(i in 1:length(file.names)){
file <- read.table(file.names[i],skip=4,header=FALSE, sep=" ", stringsAsFactors=FALSE)
st021ozone <- rbind(st021ozone, file)
}
write.table(st021ozone, file = "st021ozone",sep=",",
row.names = FALSE, qmethod = "double",fileEncoding="windows-1252")
我想加入这两个代码并制作一个完整的数据框,其中日期位于表格左侧,对应于每个文件中的 250 值。最终的数据框将有 9 行和 5 列。
这是想要的结果:
2010-02-03 250 23.54300 0.00000 12.90000
500 17.47400 3.50000 21.70000
750 41.33200 22.10000 30.40000
2010-02-10 250 36.95300 30.60000 27.10000
500 40.87700 37.80000 27.80000
750 41.46100 38.30000 30.70000
2010-02-17 250 28.91200 24.90000 5.00000
500 49.82900 32.40000 11.10000
750 50.83600 40.60000 22.20000
提前感谢您的帮助。
编辑:这是文件的一个示例。我有 53 个类似的文件:
$
### DEFAULTS ###
Color scale $
Contour levels $
Contour colors $
Contour type $
Map limits $
Map projection $
### OPTIONS ###
Image order $ 1
Title 1 left $ Case 0236-004 - Vertical
Vertical
Title 2 left $ Date and time: 2010-09-08
Receptor code: Town021
Title 3 left $
Title 4 left $
Title 1 right $
Title 2 right $ tranport
Title 3 right $ Start: 2010-01-01 00:00 UTC
Title 4 right $
gar-0236-004-01-20151103085553.png
### DATA ###
1 1 0 1 1.83400 32.00000 0.00000 21.20000
1 1 100 1 3.27300 0.00000 25.10000 21.90000
1 1 250 1 12.22200 0.00000 34.30000 25.60000
1 1 500 1 27.27400 0.00000 35.00000 31.30000
1 1 750 1 26.45300 0.00000 36.10000 35.90000
1 1 1000 1 32.62200 0.00000 36.40000 39.30000
1 1 1500 1 35.22700 0.00000 36.70000 42.20000
1 1 2000 1 37.90300 0.00000 37.40000 43.80000
1 1 3000 1 47.28200 0.00000 44.30000 46.90000
1 1 4000 1 51.01200 0.00000 49.00000 49.90000
1 1 5000 1 61.06500 0.00000 49.40000 51.00000
@alistaire 感谢您的链接。这是我尝试执行的代码。
setwd("C:/Users/")
path = "~C:/Users/"
test.sample<-""
files <- lapply(list.files(pattern = '\\.dat'), readLines)
test.sample<- rbind(test.sample, files)
do.call(rbind, lapply(files, function(lines){
# for each file, return a data.frame of the datetime, pulled with regex
data.frame(datetime = as.POSIXct(sub('^.*Date and time: ', '', lines[grep('Date and time:', lines)])),
# and the data, read in as text
read.table(text = lines[(grep('DATA', lines) + 1):length(lines)]))
}))
write.table(test.sample, file = "test.sample", sep="\t", qmethod ="double",row.names = FALSE,fileEncoding="windows-1252")
我分配了一个文件名并给出了一个路径。当我将变量名“test.sample”写入控制台时,我得到了一个矩阵:see image for the result 我想我错过了什么???注意。这里我使用的是一个包含 44 个文件的文件夹。
【问题讨论】:
标签: r dataframe data-extraction