【发布时间】:2017-08-23 20:29:22
【问题描述】:
我有一个数据框,其中包含与组 (y) 相关的数字分数,这些分数是跨不同因素 (x) 测量的,并带有结果分数。类似于下表。
BU AUDIT CORC GOV PPS TMSC TRAIN
Unit1 2.00 0.00 2.00 4.00 1.50 2.50
Unit2 3.00 1.40 3.20 1.00 1.50 3.00
Unit3 2.50 2.40 2.80 3.00 2.75 2.50
Unit4 3.00 3.20 1.60 4.00 1.00 3.00
Unit5 2.00 2.80 2.00 2.00 3.00 2.50
表是这样创建的
df %>%
group_by(BU, CC) %>% #BU = 'unit', CC = 'Control_Category
summarise(avg = mean(Score, na.rm = TRUE)) %>%
dcast(BU ~ CC, value.var = "avg") %>% print()
这些数字分数引用了一个字符串值,如下面的“表格”所示。
Control_Score > 3.499 ~ "Ineffective",
Control_Score > 2.499 & Control_Score <= 3.499 ~ "Marginally Effective",
Control_Score >= 1.500 & Control_Score <= 2.499 ~ "Generally Effective",
Control_Score > 0.000 & Control_Score <= 1.499 ~ "Highly Effective"
我尝试了一些应用函数来尝试对值进行比较。还尝试使用 case_when 变异为不可用。
最后,如果表格看起来像这样,那将是理想的:
BU, AUDIT, CORC, GOV, PPS, TMSC, TRAIN
Unit1, Generally Effective, Highly Effective, etc, etc
Unit2, Marginally Effective, Highly Effective, etc, etc
Unit3, ...,...,...
Unit4, ...,...,...
Unit5, ...,...,...
【问题讨论】:
-
0 不应该是高效的吗?第一行的第二列!
-
我在这里展示的表格只是示例,不一定对应 1:1。
-
好的。但如果他们这样做会更好。 How to make a great reproducible example in R? 这是一个很好的阅读主题。