【发布时间】:2020-06-28 16:10:58
【问题描述】:
我有一个大型数据集,其中包含多个列,每个列都有一个州名称。每行包含一个人以及他们居住在哪个州,在相应州的列中用“是”表示。
Name <- c("John", "Jane", "Joe", "Jim", "Jeane", "Jeff", "Jack")
Q1State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
Q1State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
Q1State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
Q2State1 <- c("no", "yes", "yes", "no", "no", "no", "no")
Q2State2 <- c("yes", "no", "no", "no", "no", "no", "yes")
Q2State3 <- c("no", "no", "no", "yes", "yes", "yes", "no")
DF <- data.frame(Name, Q1State1, Q1State2, Q1State3, Q2State1, Q2State2, Q2State3)
Name Q1State1 Q1State2 Q1State3 Q2State1 Q2State2 Q2State3
1 John no yes no no yes no
2 Jane yes no no yes no no
3 Joe yes no no yes no no
4 Jim no no yes no no yes
5 Jeane no no yes no no yes
6 Jeff no no yes no no yes
7 Jack no yes no no yes no
我希望以一列而不是多列来结束 State。 最终结果如下所示:
name Q1State Q2State
1 John State2 State2
2 Jane State1 State1
3 Joe State1 State1
4 Jim State3 State3
5 Jeane State3 State3
6 Jeff State3 State3
7 Jack State2 State2
我可以使用
unite(DF, State1, State2, State3)
毫无困难地完成我目标的第二部分。我的问题是所需的中间步骤。我不知道如何用适当的州名或空白填充单元格。我希望它看起来像:
name Q1State1 Q1State2 Q1State3 Q2State1 Q2State2 Q2State3
1 John State2 State2
2 Jane State1 State1
3 Joe State1 State1
4 Jim State3 State3
5 Jeane State3 State3
6 Jeff State3 State3
7 Jack State2 State2
之前发布过类似的问题Replace values in a column with specific row value from same column using loop,但该问题使用第一行数据填充单元格。我曾尝试在 dplyr 中使用类似的编码,但我无法弄清楚如何正确调用列名。
DF %>%
mutate_at(vars(starts_with('State')), ~ case_when(. == 'yes' ~colnames(.), TRUE ~ ''))
使用此代码我得到一个错误。我不确定如何指定列标题用于填充单元格。我说过尝试在 dplyr 中使用 mutate 但无法弄清楚如何正确调用列标题。
【问题讨论】:
标签: r loops dplyr columnheader