【发布时间】:2020-04-10 11:29:35
【问题描述】:
我在 R 中使用以下数据框。
输入:
structure(list(uid = c("K-1", "K-1",
"K-2", "K-3", "K-4", "K-5",
"K-6", "K-7", "K-8", "K-9",
"K-10", "K-11", "K-12", "K-13",
"K-14"), Date = c("2020-03-16 12:11:33", "2020-03-16 12:11:33",
"2020-03-16 06:13:55", "2020-03-16 10:03:43", "2020-03-16 12:37:09",
"2020-03-16 06:40:24", "2020-03-16 09:46:45", "2020-03-16 12:07:44",
"2020-03-16 14:09:51", "2020-03-16 09:19:23", "2020-03-16 09:07:37",
"2020-03-16 11:48:34", "2020-03-16 06:23:24", "2020-03-16 04:39:03",
"2020-03-16 04:59:13"), batch_no = c(7, 7, 8, 9, 9, 8,
7, 6, 7, 9, 8, 8, 7, 7, 7), marking = c("S1", "S1", "S2",
"SE_hold1", "SD_hold1", "SD_hold2", "S3", "S3", "", "SA_hold3", "S1", "S1", "S2",
"S3", "S3"), seq = c("FRD",
"FHL", NA, NA, NA, NA, NA, NA, "ABC", NA, NA, NA, NA, "DEF", NA)), .Names = c("uid",
"Date", "batch_no", "marking",
"seq"), row.names = c(NA, 15L), class = "data.frame")
uid Date batch_no marking seq
K-1 16/03/2020 12:11:33 7 S1 FRD
K-1 16/03/2020 12:11:33 7 S1 FHL
K-2 16/03/2020 12:11:33 8 SE_hold1 ABC
K-3 16/03/2020 12:11:33 9 SD_hold2 DEF
K-4 16/03/2020 12:11:33 8 S1 XYZ
K-5 16/03/2020 12:11:33 NA ABC
K-6 16/03/2020 12:11:33 7 ZZZ
K-7 16/03/2020 12:11:33 NA S2 NA
K-8 16/03/2020 12:11:33 6 S3 FRD
-
seq列将有 8 个唯一值,包括NA,不一定所有 8 个值都可用于每一天的日期。 -
batch_no将有六个唯一值,包括NA和空白,这六个值不一定适用于每一天的日期。 -
marking列将有大约 25 个唯一值,但需要将后缀为_hold#的值视为Hold,之后将有六个唯一值,包括空白和NA。
要求是按以下顺序合并 dcast 数据框,以获得用于分析的单个视图摘要。
我想在代码中保持所有唯一值不变,这样如果特定值在特定日期不可用,我将得到 0 或 - 在汇总表中。
期望的输出:
seq count percentage Marking count Percentage batch_no count Percentage
FRD 1 12.50% S1 2 25.00% 6 1 12.50%
FHL 1 12.50% S2 1 12.50% 7 2 25.00%
ABC 2 25.00% S3 1 12.50% 8 2 25.00%
DEF 1 12.50% Hold 2 25.00% 9 1 12.50%
XYZ 1 12.50% NA 1 12.50% NA 1 12.50%
ZZZ 1 12.50% (Blank) 1 12.50% (Blank) 1 12.50%
FRD 1 12.50% - - - - - -
NA 1 12.50% - - - - - -
(Blank) 0 0.00% - - - - - -
Total 8 112.50% - 8 100.00% - 8 100.00%
对于seq,我们有 % > 100,因为重复计算相同的 uid 值 FRD 和 FHL。这是公认的情况。在 Total 中将只有 uid 的不同计数。
我正在使用下面提到的 SO 代码,但无法获得所需的输出。
df = df_original %>%
mutate(marking = if_else(str_detect(marking,"hold"),"Hold", marking)) %>%
mutate_at(vars(c("seq", "batch_no", "marking")), forcats::fct_explicit_na, na_level = "(Blank)")
## You Need to do something similar with vectors of the possible values
df_combinations = purrr::cross_df(list(seq = df$seq %>% unique(),
batch_no = df$batch_no %>% unique(),
marking = df$marking %>% unique()))
df_all_combination = df_combinations %>%
left_join(df, by = c("seq", "batch_no", "marking")) %>%
group_by(seq, batch_no, marking) %>%
summarise(count = n())
【问题讨论】:
-
您能否通过分享您的数据样本来重现您的问题,以便其他人可以提供帮助(请不要使用
str()、head()或屏幕截图)?您可以使用reprex和datapasta包来帮助您。另见Help me Help you & How to make a great R reproducible example? -
@Tung:更新了有问题的
dput(datafrme)。