【发布时间】:2018-08-14 21:43:49
【问题描述】:
我有一些每周收集的数据,其中一个 sn-p 是这样的,通过dput:
p <- structure(list(railroad = structure(c(2L, 2L, 2L, 3L, 3L, 3L), .Label =
c("All Other Railroads",
"BNSF Railway Company", "CN", "CSX Transportation", "Norfolk Southern",
"The Kansas City Southern Railway and Kansas City Southern de Mexico, S.A. de
C.V. Consolidated ",
"Union Pacific Railroad"), class = "factor"), measure = structure(c(1L,
4L, 3L, 1L, 4L, 3L), .Label = c("Cars On Line - By Car Owner",
"Cars On Line - By Car Type", "Terminal Dwell (Hours)", "Train Speed (MPH)"
), class = "factor"), category = structure(c(76L, 35L, 4L, 76L,
35L, 29L), .Label = c("All Trains", "Allentown, PA", "Baltimore, MD",
"Barstow, CA", "Bellevue, OH", "Birmingham, AL", "Box", "Buffalo, NY",
"Chattanooga, TN", "Chicago (Proviso), IL", "Chicago, IL", "Cincinnati, OH",
"Coal Unit", "Columbus, OH", "Conway, PA", "Corbin, KY", "Covered Hopper",
"Decatur, IL", "Denver, CO", "Elkhart, IN", "Entire Railroad",
"Fond du Lac Yard, WI", "Foreign RR", "Fort Worth, TX", "Galesburg, IL",
"Gondola", "Grain Unit", "Hamlet, NC", "Harrison Yard (Memphis), TN",
"Hinkle, OR", "Houston (Englewood), TX", "Houston (Settegast), TX",
"Houston, TX", "Indianapolis, IN", "Intermodal", "Jackson Yard, MS",
"Jackson, MS", "Kansas City, KS", "Kansas City, MO", "Knoxville, TN",
"Laredo, TX", "Lincoln, NE", "Linwood, NC", "Livonia, LA", "Louisville, KY",
"MacMillan Yard (Toronto), ON", "Macon, GA", "Manifest", "Markham Yard, IL",
"Memphis, TN", "Monterrey, NL", "Montgomery, AL", "Multilevel",
"Nashville, TN", "New Orleans, LA", "North Little Rock, AR",
"North Platte East, NE", "North Platte West, NE", "Northtown, MN",
"Nuevo Laredo, TM", "Open Hopper", "Other", "Pasco, WA", "Pct. Private",
"Pine Bluff, AR", "Private", "Roanoke, VA", "Roseville, CA",
"Russell, KY", "San Luis Potosi, SL", "Sanchez, TM", "Selkirk, NY",
"Sheffield, AL", "Shreveport, LA", "Symington Yard (Winnipeg), MB",
"System", "Tank", "Tascherau Yard (Montreal), QC", "Thornton Yard (Vancouver),
BC",
"Toledo, OH", "Total", "Tulsa, OK", "Walker Yard (Edmonton), AB",
"Waycross, GA", "West Colton, CA", "Willard, OH"), class = "factor"),
`201510` = c(66923, 33.9, 39.3, 40227, 30.8, 17.5), `201510` = c(66637,
32.6, 56.6, 40778, 30.9, 18.3), `201510` = c(66309, 33.4,
44.9, 40407, 30.5, 17.3), `201511` = c(65980, 34.6, 37.5,
40316, 30.6, 17.5), `201511` = c(67034, 34.6, 43.1, 40174,
30.4, 18.7)), row.names = c(1L, 15L, 21L, 33L, 47L, 53L), class =
"data.frame")
共有 143 列,第 4 - 143 列是数字。我想计算具有相同列名的所有列的平均值。所以下面有列201510重复3次和列201511重复两次。所需的输出是重复的每列的平均值。例如,201510 将具有以下值:
`201510`
[1] 66623.00000 33.30000 46.93333 40470.66667 30.73333 17.70000
我尝试了以下代码:
library(tidyverse)
p = data.frame(p)
p %>%
gather(time,value,railroad, measure, category) %>%
mutate(time = gsub('X([^.]+)|.', '\\1', time)) %>%
group_by(time, value, railroad, measure, category) %>%
summarise(MEAN = mean(value)) %>%
ungroup() %>%
spread(time, MEAN)
这会产生以下错误:
`Error in grouped_df_impl(data, unname(vars), drop) :
Column `railroad` is unknown
In addition: Warning message:
attributes are not identical across measure variables;
they will be dropped `
有没有办法做到这一点?
【问题讨论】:
-
到目前为止你尝试了什么?
-
@camille 刚刚编辑了代码