【问题标题】:dplyr mutate: Excluding observations similar to the current onedplyr mutate:排除与当前相似的观察
【发布时间】:2018-08-27 13:17:21
【问题描述】:

我有一些这样的数据:

X   Y
-----
A   1
A   2
B   3
B   4
C   5
C   6

我想添加一个新列,其值等于行中所有 Y 的平均值,其中 X 与当前观察的 X 不相等。 在这种特殊情况下,我们会得到

X   Y   Mean
-------------------
A   1   (3+4+5+6)/4
A   2   (3+4+5+6)/4
B   3   (1+2+5+6)/4
B   4   (1+2+5+6)/4
C   5   (1+2+3+4)/4
C   6   (1+2+3+4)/4

提前致谢!

【问题讨论】:

    标签: r mean dplyr


    【解决方案1】:

    您可以更简洁地执行此操作,但这会给您带来结果。

    您实际上创建了一个列,其中包含整个 data.frame 的总观察值和记录总和。然后按X 列分组并重复该过程,通过取差值可以计算平均值。

    数据

    df <- data.frame(X = c("A", "A", "B", "B", "C", "C"),
                     Y = c(1:6))
    

    解决方案

    library(tidyverse)
    df %>%
      mutate(total_sum = sum(Y),
             total_obs = n()) %>%
      group_by(X) %>%
      mutate(group_sum = sum(Y),
             group_obs = n()) %>%
      ungroup() %>%
      mutate(other_group_sum = total_sum - group_sum,
             other_group_obs = total_obs - group_obs,
             other_mean = other_group_sum/other_group_obs) %>%
      select(X, Y, other_mean)
    

    结果

    # A tibble: 6 x 3
      X         Y other_mean
      <fct> <int>      <dbl>
    1 A         1       4.50
    2 A         2       4.50
    3 B         3       3.50
    4 B         4       3.50
    5 C         5       2.50
    6 C         6       2.50
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-06-21
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-05-13
      • 1970-01-01
      相关资源
      最近更新 更多