【发布时间】:2020-08-29 07:15:32
【问题描述】:
下面的代码创建的原始数据类似于我正在使用的数据。我使用 tibble 包中的 add_row 函数编写了一些代码来重新格式化它。现在我收到一个错误(此代码在 2020 年 4 月之前有效)。看起来由于包的更新,子集的规则变得更严格了?我想知道是否有人可以帮助纠正这个错误...... 首先创建数据
# Create replicate of raw data
date <- seq(from = as.Date('1999-01-01'),
to = as.Date('2013-12-31'),
by = 'day')
temp <- rnorm(5479,15,5)
precip <- rlnorm(5479)
rawdata <- data.frame(date=date,
temp=round(temp, digits = 2),
precip=round(precip, digits = 2))
# Add columns needed to run code
rawdata$year <- as.numeric(substr(rawdata$date,1,4))
rawdata$month <- as.numeric(substr(rawdata$date, 6,7))
rawdata$chardate <- format(rawdata$date, '%Y-%h-%d') # create abbreviated month column
rawdata$charmonth <- substr(rawdata$chardate, 6,8) # for formatting
rawdata$charmonth <- as.character(rawdata$charmonth)
rawdata$day <- as.numeric(substr(rawdata$date, 9,10))
rawdata$uniqdate <- rawdata$year*100+as.numeric(rawdata$day)+rawdata$month*10
rawdata$uniqmonth <- (rawdata$year*100)+rawdata$month# create unique month identifier
rawdata$yr <- NA # This column will be filled only in the new rows to be added
# Create weather object to feed the for loop below----
weather <- data.frame(year = rawdata$year,
month = rawdata$month,
day = rawdata$day,
charmonth = rawdata$charmonth,
uniqmonth = rawdata$uniqmonth,
uniqdate = rawdata$uniqdate,
temp = rawdata$temp,
precip = rawdata$precip,
yr = rawdata$yr)
# weather$charmonth <- as.character(rawdata$charmonth)
现在出现错误...我正在尝试在每个月的顶部添加一行数据,其中包含该月的天数,缩写为三个字母的月份(jan、feb、mar 等。 ) 和年份。
library(tibble) # package containing the add_row function
# create empty list to put all of the monthly dataframes in
newdat <- list()
# the following loop will create a dataframe for each month and put in a list
for(i in unique(weather$uniqmonth)) { # for every unique month value
# create object nam that is of the format 'df.uniqmonth'
nam <- paste("df", i, sep = ".")
# create object dat that contains all data for each unique month
dat <- weather[weather$uniqmonth==i,]
# add a row of data at the start of each dataframe with the days in month, month abbr., year
dat <- add_row(dat, year = NA, month = NA, day = NA,
charmonth = NA, uniqmonth = NA, uniqdate = NA,
# the line below is the info we are adding in the columns we will keep
temp = na.omit(max(dat$day)), precip = unique(dat$charmonth), yr = unique(dat$year),
.before = 1)
# just keep required columns
dat <- data.frame(dat$temp, dat$precip, dat$yr)
# add new dataframe to a list, using the new name
newdat[[nam]] <- dat
}
**你可以运行循环或者只是逐行(设置i = 199901),错误是一样的:
错误:无法组合 ..1$precip 和 ..2$precip。
最终我应该能够运行以下命令来获得我需要的输出,并在文本编辑器中完成(删除尾随逗号)。**
# Merge all data into a dataframe
full_data <- do.call("rbind", newdat)
# turn NA's into blanks
full_data[is.na(full_data)] <- ""
**感谢任何帮助解决此错误!
这是我需要的最终产品
a <- c("Jan",
round(rnorm(31,15,5), digits = 2),
"Feb",
round(rnorm(28,5,5), digits = 2),
"Mar",
round(rnorm(31,15,5),digits = 2))
b <- c(31,
rlnorm(31),
28,
rlnorm(28),
31,
rlnorm(31))
c <- c(1999,
rep(NA,31),
1999,
rep(NA,28),
1999,
rep(NA,31))
final_data <- data.frame(temp = a,
precip = round(b,digits=2),
year = c)
【问题讨论】:
-
我不明白为什么用
precip=unique(dat$charmonth)(character类)添加一行可以附加到dat$precip类numeric。 -
坦率地说,当
i <- 1时,我得到一个不同的(但同样可以避免的)错误。 -
我开始研究这个,但我担心我推断 waaaaay 太多了。我不明白你为什么要将
max(dat$day)分配给temp,将unique(date$charmonth)分配给precip。如果这是故意的,那么......除了建议您将所有内容转换为字符串并希望数据的下游使用能够适应之外,我不知道如何帮助您。 (不,这不是一个真正的建议,但目前的代码对我来说毫无意义。) -
此外,您在其中使用
unique添加“一行”表明您打算生成n行。这假设您的所有unique调用将产生相同的长度(或1),否则您将得到回收(这是一个数据损坏问题)或错误(这应该真的是非 1 回收的默认行为)。我怀疑这一切都可以通过一个简单的weather %>% group_by(uniqmonth) %>% summarize(...) %>% bind_rows(weather)(而不是for循环和稍后rbinding)来完成,但是代码的问题让我很困惑。 -
(最后,您的
full_data[is.na(full_data)] <- ""确实在执行我在两个cmets 前开玩笑地建议的事情:它将所有字段(至少包含一个NA)转换为character。)