【问题标题】:Filling row elements with column heading given certain criteria in R给定R中的某些条件,用列标题填充行元素
【发布时间】:2020-06-28 16:10:58
【问题描述】:

我有一个大型数据集,其中包含多个列,每个列都有一个州名称。每行包含一个人以及他们居住在哪个州,在相应州的列中用“是”表示。

Name <- c("John", "Jane", "Joe", "Jim", "Jeane", "Jeff", "Jack")
Q1State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
Q1State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
Q1State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
Q2State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
Q2State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
Q2State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
DF <- data.frame(Name, Q1State1, Q1State2, Q1State3, Q2State1, Q2State2, Q2State3)

   Name Q1State1 Q1State2 Q1State3 Q2State1 Q2State2 Q2State3
1  John       no      yes       no       no      yes       no
2  Jane      yes       no       no      yes       no       no
3   Joe      yes       no       no      yes       no       no
4   Jim       no       no      yes       no       no      yes
5 Jeane       no       no      yes       no       no      yes
6  Jeff       no       no      yes       no       no      yes
7  Jack       no      yes       no       no      yes       no

我希望以一列而不是多列来结束 State。 最终结果如下所示:

   name    Q1State   Q2State
1  John     State2    State2
2  Jane     State1    State1
3   Joe     State1    State1
4   Jim     State3    State3
5 Jeane     State3    State3
6  Jeff     State3    State3
7  Jack     State2    State2

我可以使用 unite(DF, State1, State2, State3) 毫无困难地完成我目标的第二部分。我的问题是所需的中间步骤。我不知道如何用适当的州名或空白填充单元格。我希望它看起来像:

   name Q1State1 Q1State2 Q1State3  Q2State1  Q2State2  Q2State3
1  John            State2                       State2          
2  Jane State1                        State1                        
3   Joe State1                        State1                    
4   Jim                     State3                        State3
5 Jeane                     State3                        State3
6  Jeff                     State3                        State3 
7  Jack            State2                       State2   

之前发布过类似的问题Replace values in a column with specific row value from same column using loop,但该问题使用第一行数据填充单元格。我曾尝试在 dplyr 中使用类似的编码,但我无法弄清楚如何正确调用列名。

DF %>% 
  mutate_at(vars(starts_with('State')), ~ case_when(. == 'yes' ~colnames(.), TRUE ~ ''))

使用此代码我得到一个错误。我不确定如何指定列标题用于填充单元格。我说过尝试在 dplyr 中使用 mutate 但无法弄清楚如何正确调用列标题。

【问题讨论】:

    标签: r loops dplyr columnheader


    【解决方案1】:

    一种可能是:

    DF %>%
     transmute(Name,
               State = names(.)[max.col(. == "yes")])
    
       Name  State
    1  John State2
    2  Jane State1
    3   Joe State1
    4   Jim State3
    5 Jeane State3
    6  Jeff State3
    7  Jack State2
    

    更新问题的选项,添加tidyr:

    DF %>%
     pivot_longer(-Name) %>%
     extract(name, into = c("name1", "name2"), "(Q*\\d+)([[:alnum:]]+)") %>%
     filter(value == "yes") %>%
     select(-value) %>%
     mutate(name1 = paste0(name1, "State")) %>%
     pivot_wider(names_from = "name1", values_from = "name2") 
    
      Name  Q1State Q2State
      <chr> <chr>   <chr>  
    1 John  State2  State2 
    2 Jane  State1  State1 
    3 Joe   State1  State1 
    4 Jim   State3  State3 
    5 Jeane State3  State3 
    6 Jeff  State3  State3 
    7 Jack  State2  State2 
    

    【讨论】:

    • 这可行,但有没有办法在调用名称时合并starts_with?我需要生成对应于不同问题的多个状态列。对于每个问题,我有 50 列,其标题是问题编号和状态(即 Q1 State1、Q2 State1、Q3 State1)。我想在生成新列时指定我想要的问题。我可以使用 select 创建多个数据框,然后将它们重新连接在一起。但我宁愿把所有东西都合二为一。
    • 请添加一些示例数据和您想要的输出。没有它,很难看到你在找什么:)
    • 我已经编辑了我的原始帖子,希望能更好地代表我希望完成的工作。出于隐私原因,我真的不能给你我的数据集样本。但是,编辑后的代码与我的数据设置相匹配,除了我有 State1-50 和 Q1-5 的每个问题编号。
    • 更新了我的帖子:)
    【解决方案2】:

    您可以转换为长格式并过滤:

    Name <- c("John", "Jane", "Joe", "Jim", "Jeane", "Jeff", "Jack")
    State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
    State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
    State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
    DF <- data.frame(Name, State1, State2, State3)
    
    DF %>%
      pivot_longer(-Name, names_to = "State", values_to = "value") %>%
      filter(value == "yes") #%>%
      # select(-value)
    
    # # A tibble: 7 x 3
    # Name  State  value
    # <fct> <chr>  <fct>
    #   1 John  State2 yes  
    # 2 Jane  State1 yes  
    # 3 Joe   State1 yes  
    # 4 Jim   State3 yes  
    # 5 Jeane State3 yes  
    # 6 Jeff  State3 yes  
    # 7 Jack  State2 yes  
    

    【讨论】:

      【解决方案3】:

      data.table 的选项

      library(data.table)
      melt(setDT(DF), id.var = 'Name', variable.name = 'State')[
               value == 'yes'][, value := NULL][]
      #    Name  State
      #1:  Jane State1
      #2:   Joe State1
      #3:  John State2
      #4:  Jack State2
      #5:   Jim State3
      #6: Jeane State3
      #7:  Jeff State3
      

      数据

      Name <- c("John", "Jane", "Joe", "Jim", "Jeane", "Jeff", "Jack")
      State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
      State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
      State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
      DF <- data.frame(Name, State1, State2, State3)
      

      【讨论】:

        【解决方案4】:

        使用 data.table

        DF2 <- dcast(melt(DF, id.vars="Name")[value == "yes"][, c("Q", "State") := tstrsplit(variable, "State")][, -c("value", "variable")], ... ~ Q)
        

        给予

            Name Q1 Q2
        1:  Jack  2  2
        2:  Jane  1  1
        3: Jeane  3  3
        4:  Jeff  3  3
        5:   Jim  3  3
        6:   Joe  1  1
        7:  John  2  2
        

        【讨论】:

          猜你喜欢
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 1970-01-01
          • 2020-09-28
          • 1970-01-01
          相关资源
          最近更新 更多