【问题标题】:Best way to input daily data into R to allow further manipulation将每日数据输入 R 以允许进一步操作的最佳方式
【发布时间】:2016-03-10 12:41:54
【问题描述】:

我在 Excel 中有每日降雨数据(我可以将其保存为 CSV 或 txt 文件),我想对其进行操作并加载到 R 中。我对 R 很陌生。

数据的格式是这样的,我有以下列

年份;月;每月第 1 天下雨,第 2 天下雨,...,第 31 天下雨;

这意味着我有一个大数组/表。有些数据丢失是因为没有记录,有些是因为 2 月 31 日、6 月 31 日等不存在。

我想分析每月总计及其分布等内容。

输入数据的最佳方式是什么,以便可以轻松操作,并且我可以区分缺失数据和 NULL 数据(2 月 31 日)?

提前非常感谢

【问题讨论】:

  • 旁注:提供您尝试过的示例数据和代码总是一个好主意,人们可以使用它们 - 这将为您提供更多答案,并且您可能会更快地收到它们。否则,每个社区成员都必须为自己/自己建立一个榜样。见stackoverflow.com/questions/5963269/…

标签: r import


【解决方案1】:

有几件事供您查看。例如。 readxl::read_excel() 用于读取 excel 文件或Hmisc::monthDays(dates) 用于确定日期向量中每个月的天数。

无论如何,这里有一个作为入门的想法:

# create sample data
set.seed(1)
mat <- matrix(rbinom(5*31, 31, .5), nrow=5)
mat[sample(1:length(mat), 10)] <- NA
df <- data.frame(year=2016, month=1:5, mat)

#  reshape data from wide to long format
library(reshape2)
dflong <- melt(df, id.vars = 1:2, variable.name = "day")

# add date column (will be NA if conversion is not possible, i.e. if date does not exists)
dflong$date <- as.Date(with(dflong, paste(year, month, day, sep="-")), format = "%Y-%m-X%e")

# Select only existing dates
dflong <- subset(dflong[order(dflong$month), ], !is.na(date))

# Aggregate: means per month and year (missing values removed)
aggregate(value~year+month, dflong, mean, na.rm=TRUE)
#   year month    value
# 1 2016     1 15.93548
# 2 2016     2 15.26923
# 3 2016     3 15.10345
# 4 2016     4 15.74074
# 5 2016     5 16.16667

【讨论】:

  • 谢谢,很抱歉我的回复慢了。我会研究你可能的解决方案,看看它是否能让我继续前进。干杯
猜你喜欢
  • 2021-01-05
  • 2021-12-25
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多