【发布时间】:2020-01-31 20:01:39
【问题描述】:
我在下面有两个数据框:
输入输出df1:
structure(list(Location = c("1100 2ND AVENUE", "1100 2ND AVENUE",
"1100 2ND AVENUE", "1100 2ND AVENUE", "1100 2ND AVENUE", "1100 2ND AVENUE"
), `Ivend Name` = c("3 Mskt 1.92oz", "Almond Joy 1.61oz", "Aquafina 20oz",
"BCanyonChptleAdzuk1.5oz", "BlkForest FrtSnk 2.25oz", "BluDimndSmkhseAlmd1.5oz"
), `Category Name` = c("Candy", "Candy", "Water", "Salty Snacks",
"Candy", "Nuts/Trailmix"), Calories = c(240, 220, 0, 215, 193,
260), Sugars = c("36", "20", "0", "2", "32", "2"), Month = structure(c(4L,
4L, 4L, 4L, 4L, 4L), .Label = c("Oct", "Nov", "Dec", "Jan", "Feb",
"Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep"), class = "factor"),
Products_available_per_machine = c(0, 0, 0, 0, 0, 0), Units_sold = c(0,
0, 0, 0, 0, 0), Total_Sales = c(0, 0, 0, 0, 0, 0), Spoils = c(0,
0, 0, 0, 0, 0), Building = c("1100 2ND", "1100 2ND", "1100 2ND",
"1100 2ND", "1100 2ND", "1100 2ND"), Item = structure(c(2L,
2L, 1L, 2L, 2L, 2L), .Label = c("Beverage", "Food"), class = "factor"),
Year = structure(c(1L, 1L, 1L, 1L, 1L, 1L), .Label = "2019", class = "factor")), row.names = c(NA,
-6L), class = c("data.table", "data.frame"), .internal.selfref = <pointer: 0x00000233561b1ef0>)
输入输出df2:
structure(list(`Date Ran` = structure(c(1548892800, 1551312000,
1553817600, 1556582400, 1561680000, 1564531200), class = c("POSIXct",
"POSIXt"), tzone = "UTC"), Year = c(2019, 2019, 2019, 2019, 2019,
2019), Month = c("January", "February", "March", "April", "June",
"July"), Location = c("SEA18", "SEA18", "SEA18", "SEA18", "SEA18",
"SEA18"), Building = c("Alexandria", "Alexandria", "Alexandria",
"Alexandria", "Alexandria", "Alexandria"), Population = c(1177,
1179, 1178, 1156, 1163, 1163)), row.names = c(NA, -6L), class = c("tbl_df",
"tbl", "data.frame"))
我想从 DF 2 中提取 pop col 并将其添加到基于“Building”和“Month”的 Dataframe 1,以便填充 DF2 中的人口。
我使用合并尝试了此命令,但执行时 col 为 NULL:
df_2019_final1$Population <- df_2019_pop$Population[match(df_2019_final1$Month, df_2019_pop$Month, df_2019_final1$Building, df_2019_pop$Building)]
subset_df_pop <- df_2019_pop[, c("Month", "Building", "Population")]
updated_2019_test <- merge(df_2019_final1, subset_df_pop, by = c('Month', 'Building'))
两者都产生 NULLS 和空白 DF
任何帮助将不胜感激。
【问题讨论】:
-
match只需要 2 个向量match(x, table, nomatch = NA_integer_, incomparables = NULL) -
你可以试试
merge(df1, df2, by = c('Month', 'Building')) -
我刚尝试合并,它输出了一个空白的df。出于某种原因,它把所有的列都带了过来,而不是 Month/Building
-
您需要通过仅包含这些列 + Populatoin 即
merge(df1, df2[, c("Month", "Building", "Population")], by = c('Month', 'Building'))来细分第二个数据 -
我用你的解决方案编辑了我的问题,但之后仍然得到一个空白的 df。