【问题标题】:replace NA with values from another table based on groupings (not one-by-one lookup table)将 NA 替换为基于分组的另一个表中的值(不是一对一的查找表)
【发布时间】:2017-01-20 07:18:10
【问题描述】:

我的目标是用另一个查找表中的值替换一个表中的值。有一个问题:此查找表不是Replace na's with value from another df 中讨论的逐一查找表,但查找将基于多个列分组完成。因此,如果根据查找表中的这些分组返回多个条目,则所有这些条目都需要填充到原始表中。

我能够完成这项任务,但我需要两件事的帮助:

a) 我的代码真的很乱。每次我必须做类似的事情时,我最终都会花费大量时间试图弄清楚我做了什么,然后重新使用它。所以,我会很感激任何更干净、更简单的东西。

b) 速度很慢。我有多个 ifelse 声明。当我在 36M 记录的实际数据上运行此程序时,会花费很多时间。

这是我的虚拟数据来源:

dput(DFile)
structure(list(Region_SL = c("G1", "G1", "G1", "G1", "G2", "G2", 
"G3", "G3", "G3", "G3", "G4", "G4", "G4", "G4", "G5", "G5"), 
    Country_SV = c("United States", "United States", "United States", 
    "United States", "United States", "United States", "United States", 
    "United States", "United States", "United States", "United States", 
    "United States", "United States", "United States", "UK", 
    "UK"), Product_BU = c("Laptop", "Laptop", "Laptop", "Laptop", 
    "Laptop", "Laptop", "Laptop", "Laptop", "Laptop", "Laptop", 
    "Laptop", "Laptop", "Laptop", "Laptop", "Power Cord", "Laptop"
    ), Prob_model3 = c(0, 79647405.9878251, 282615405.328728, 
    NA, NA, 363419594.065383, 0, 72870592.8458704, 260045174.088548, 
    369512727.253779, 0, 79906001.2878251, 285128278.558728, 
    405490639.873629, 234, NA), DoS.FY = c(2014, 2013, 2012, 
    NA, 2015, 2015, 2015, 2015, 2015, 2015, 2015, 2015, 2015, 
    2015, 2016, NA), Insured = c("Covered", "Covered", "Covered", 
    NA, NA, "Not Covered", "Not Covered", "Not Covered", "Not Covered", 
    "Not Covered", "Not Covered", "Not Covered", "Not Covered", 
    "Not Covered", "Covered", NA)), .Names = c("Region_SL", "Country_SV", 
"Product_BU", "Prob_model3", "DoS.FY", "Insured"), row.names = c(NA, 
16L), class = "data.frame")

这是我的分组查找表:

dput(Master_Joined)
structure(list(Region_SL = c("G1", "G1", "G1", "G1", "G2", "G3", 
"G4", "G5", "G5", "G5"), Country_SV = c("United States", "United States", 
"United States", "United States", "United States", "United States", 
"United States", "UK", "UK", "UK"), Product_BU = c("Laptop", 
"Laptop", "Laptop", "Laptop", "Laptop", "Laptop", "Laptop", "Power Cord", 
"Laptop", "Laptop"), DoS.FY = c(2014, 2013, 2012, 2015, 2015, 
2015, 2015, 2016, 2017, 2017), Insured = c("Covered", "Covered", 
"Covered", "Uncovered", "Not Covered", "Not Covered", "Not Covered", 
"Covered", "Uncovered", "Covered")), .Names = c("Region_SL", 
"Country_SV", "Product_BU", "DoS.FY", "Insured"), row.names = c(NA, 
10L), class = "data.frame")

从某种意义上说,这是“分组”的,所有条目都是唯一的。

最后,这是我的代码:

#Which fields are missing?
Missing<-DFile[is.na(DFile$Prob_model3),]

Column_name<-colnames(DFile)[4]
colnames(DFile)[4]<-"temp_prob"

#Replace Prob_model3
DFile<-DFile %>%
  group_by(Region_SL, Country_SV, Product_BU) %>%
  dplyr::mutate(Average_Value = mean(temp_prob,na.rm = TRUE)) %>%
  rowwise() %>%
  dplyr::mutate(Col_name1 = ifelse(is.na(temp_prob),Average_Value,temp_prob)) %>%
  dplyr::select(Region_SL:Product_BU,DoS.FY,Insured,Col_name1)

colnames(DFile)[6]<-Column_name

  Missing$DoS.FY<-NULL

  Missing_FYear<-Missing %>% 
    inner_join(Master_Joined,by = c("Region_SL", "Country_SV", "Product_BU")) %>%
    group_by(Region_SL, Country_SV, Product_BU, DoS.FY, Insured.y) %>%
    dplyr::distinct() %>%
    left_join(Missing)

  Missing_FYear$Prob_model3<-NULL

  DFile <-DFile %>% 
    left_join(Missing_FYear,by = c("Region_SL", "Country_SV", "Product_BU", "Insured")) %>%
    dplyr::rowwise() %>%
    mutate(DoS.FY=ifelse((is.na(`DoS.FY.y`)|is.na(`DoS.FY.x`)),sum(`DoS.FY.y`,`DoS.FY.x`,na.rm=TRUE),`DoS.FY.x`), Insured_Combined = ifelse(is.na(Insured),Insured.y,Insured)) %>%
    dplyr::select(Region_SL:Product_BU,Prob_model3,DoS.FY, Insured_Combined)  

  colnames(DFile)[6]<-"Insured"
  #Check again
  Missing<-DFile[is.na(DFile$Prob_model3),] 

  if (nrow(Missing) > 1)
  { #you have NaNs, replace them with 0
    DFile[is.nan(DFile$Prob_model3),"Prob_model3"] <- 0
   }
  Missing<-DFile[is.na(DFile$Prob_model3),] 

预期输出:DFile 与运行上述代码后一样。

非常感谢您的帮助。我已经为这个问题苦苦挣扎了大约一个星期。

【问题讨论】:

  • @Sotos 和 Jonathan - 感谢您的指出。 myout 与通过代码运行 DFile 后完全相同。如果你愿意,我可以再发myout;我删除了它,因为我不确定出了什么问题。我在运行代码后通过复制粘贴DFile 生成了它。不确定这是否导致了问题。

标签: r dplyr


【解决方案1】:

一个想法是找到具有NA 的Region_SL。一旦我们这样做了,我们使用plyr 的rbind.fill 来绑定到new_df。然后,我们过滤掉所有带有NA 的行(最后一列除外 - 第 6 列)。我们创建一个新变量Prob_model4,它保存Region_SL 的每组均值。然后我们使用coalesce 来“合并”这两列。

library(dplyr)
ind <- unique(as.integer(which(is.na(DFile), arr.ind = TRUE)[,1]))
new_df <- plyr::rbind.fill(Master_joined[Master_joined$Region_SL %in% DFile$Region_SL[ind],], DFile)

new_df %>% 
  arrange(Region_SL, Prob_model3) %>% 
  filter(complete.cases(.[-6])) %>% 
  group_by(Region_SL) %>% 
  mutate(Prob_model3 = replace(Prob_model3, is.na(Prob_model3), mean(Prob_model3, na.rm = T))) %>%  
  ungroup()

# A tibble: 21 × 6
#   Region_SL    Country_SV Product_BU DoS.FY     Insured Prob_model3
#       <chr>         <chr>      <chr>  <dbl>       <chr>       <dbl>
#1         G1 United States     Laptop   2014     Covered           0
#2         G1 United States     Laptop   2013     Covered    79647406
#3         G1 United States     Laptop   2012     Covered   282615405
#4         G1 United States     Laptop   2014     Covered   120754270
#5         G1 United States     Laptop   2013     Covered   120754270
#6         G1 United States     Laptop   2012     Covered   120754270
#7         G1 United States     Laptop   2015   Uncovered   120754270
#8         G2 United States     Laptop   2015 Not Covered   363419594
#9         G2 United States     Laptop   2015 Not Covered   363419594
#10        G3 United States     Laptop   2015 Not Covered           0
# ... with 11 more rows

【讨论】:

  • 当我运行问题中的命令时,我得到了 20 行的框架。你的答案生成一个 21?
  • @JonathanvonSchroeder 是的,我注意到了。如果我删除重复项,我会得到一个包含 19 行的数据框,这是正确的,因为 OP 的数据框有 20 行,其中 1 行重复(G2 - 第 9 行)。我选择这样保留它,以便 OP 可以随心所欲地进行操作
【解决方案2】:

另一种思考方式是仅将 DoS.FY 或 Insured 中缺失值的行与主数据合并:

#Replace missing probabilities by grouped average
DFile_new <- DFile %>% group_by(Region_SL,Country_SV,Product_BU) %>% mutate(Prob_model3 = coalesce(Prob_model3,mean(Prob_model3, na.rm = T))) %>% ungroup()
#This leads to one NaN because for
# 16        G5            UK     Laptop          NA     NA        <NA>
#there are no other rows in the same group
DFile_new$Prob_model3[is.nan(DFile_new$Prob_model3)] <- 0

#Split dataset into two parts
#1) The part that has no NA's in DoS.FY and Insured
DFile_new1 <- filter(DFile_new,!is.na(DoS.FY) & !is.na(Insured))
#2) The part has NA's in either DoS.FY or Insured
DFile_new2 <- filter(DFile_new,is.na(DoS.FY) | is.na(Insured))

#merge DFile_new2 and Master_Joined
DFile_new2 <- merge(DFile_new2,Master_Joined,by=c("Region_SL","Country_SV","Product_BU")) %>%
  mutate(DoS.FY.x = coalesce(DoS.FY.x,DoS.FY.y), Insured.x = coalesce(Insured.x,Insured.y)) %>%
  select(-Insured.y,-DoS.FY.y) %>% rename(Insured=Insured.x, DoS.FY = DoS.FY.x)

#Put all rows in frame
my_out_new <- rbind(DFile_new1,DFile_new2)

这会产生与 OP 的代码相同的结果(尽管顺序不同):

> compare <- function(df1,df2) {
+   idx1 <- c()
+   idx2 <- c()
+   for(i in 1:nrow(df1)) {
+     found <- FALSE
+     for(j in 1:nrow(df2)) {
+       if(!(j %in% idx2)) {
+         idx = as.logical(df1[i,] != df2[j,])
+         d <- suppressWarnings(abs(as.numeric(df1[i,idx])-as.numeric(df2[j,idx]))) < 1e-5
+         if(!(any(is.na(d))) & all(d)) {
+           idx1 <- c(idx1,i)
+           idx2 <- c(idx2,j)
+           break;
+         }
+       }
+     }
+   }
+   rbind(idx1,idx2)
+ }
> compare(my_out,my_out_new)
     [,1] [,2] [,3] [,4] [,5] [,6] [,7] [,8] [,9] [,10] [,11] [,12] [,13] [,14] [,15] [,16] [,17]
idx1    1    2    3    4    5    6    7    8    9    10    11    12    13    14    15    16    17
idx2    1    2    3   14   15   16   17    4   18     5     6     7     8     9    10    11    12
     [,18] [,19] [,20]
idx1    18    19    20
idx2    13    19    20

(其中 my_out 是 OP 代码的结果 DFile)

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2013-08-06
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-08-13
    • 1970-01-01
    • 2016-11-25
    • 1970-01-01
    相关资源
    最近更新 更多