【问题标题】:Reshape dataframe so that each entry of column in repeated over all other columns重塑数据框,以便列的每个条目在所有其他列上重复
【发布时间】:2019-04-16 14:10:50
【问题描述】:

我有一个data.frame,dat,看起来像这样

dat = data.frame(x = c(1, 1.1, 1.2, 1.3), y = c(2, 2.1, 2.2, 2.3), output = c(2, 10, 101, 100))
    x   y output
1 1.0 2.0      2
2 1.1 2.1     10
3 1.2 2.2    101
4 1.3 2.3    100

我希望列“x”和“输出”的每一对元素在“y”列上重复。

我曾尝试使用tidyr::spread、tidyr::gather 和reshape2::melt 无济于事。这是因为我是使用tidyr 和reshape2 以及其他整形包的初学者。

目前,我使用循环从列“x”和“输出”中提取每个元素对,并创建一个新的 data.frame final_df,它结合了生成的 data.frames。我相信这绝对不是最有效的方法,并且我相信某处有一个单行函数可以为我发挥这种魔力。

在生成的 data.frame 中,如果我使用 say 对 data.frame 进行子集化,

dplyr::filter(final_df, x == 1, output == 2)

应该是这样的:

data.frame(x = rep(1, dat$x[1], nrow(dat)), y = dat$y, output = rep(dat$output[1], nrow(dat)))
  x   y output
1 1 2.0      2
2 1 2.1      2
3 1 2.2      2
4 1 2.3      2

我会对使用 tidyverse 的答案感到满意。谢谢。

【问题讨论】:

  • 您能否解释一下这部分:each pair of elements of columns "x" and "output" is repeated over column "y"?您希望一对重复多少次?
  • 如果有意义的话,我想保持 y 固定并重复每个 x 和输出对以匹配 y 的长度。

标签: r tidyverse tidyr reshape2


【解决方案1】:

这是一种选择

library(dplyr)
library(tidyr)
dat %>% mutate(y1=paste(y,collapse = ',')) %>% separate_rows(y1)

如果 x 和 output 中没有重复,即我们可以将它们视为 ID 列,那么我们可以使用tidyr::complete

dat %>% complete(nesting(x,output),y)

【讨论】:

    【解决方案2】:

    一种解决方案:

    require(dplyr)
    require(tidyr)
     dat %>% select(-y) %>% crossing(dat %>% select(y))
    
         x output   y
    1  1.0      2 2.0
    2  1.0      2 2.1
    3  1.0      2 2.2
    4  1.0      2 2.3
    5  1.1     10 2.0
    6  1.1     10 2.1
    7  1.1     10 2.2
    8  1.1     10 2.3
    9  1.2    101 2.0
    10 1.2    101 2.1
    11 1.2    101 2.2
    12 1.2    101 2.3
    13 1.3    100 2.0
    14 1.3    100 2.1
    15 1.3    100 2.2
    16 1.3    100 2.3
    

    【讨论】:

      猜你喜欢
      • 2018-09-11
      • 1970-01-01
      • 2020-08-18
      • 1970-01-01
      • 2018-07-15
      • 1970-01-01
      • 2013-04-25
      • 2021-07-23
      • 1970-01-01
      相关资源
      最近更新 更多