【问题标题】:tidyr::spread() function throws an errortidyr::spread() 函数抛出错误
【发布时间】:2017-06-03 23:41:12
【问题描述】:

我尝试在tidyverse包中使用gather和spread函数,但是它在spread函数中抛出错误

库(插入符号)

dataset<-iris

# gather function is to convert wide data to long data

dataset_gather<-dataset %>% tidyr::gather(key=Type,value = Values,1:4)

head(dataset_gather)

# spead is the opposite of gather

下面的代码会抛出类似这样的错误 Error: Duplicate identifiers for rows

dataset_spead<- dataset_gather%>%tidyr::spread(key = Type,value = Values)

【问题讨论】:

  • 使用gather后你没有唯一标识符,所以你不能在这样的数据帧上使用spread
  • 也许如果您包含一些您可以使用的实际数据以及您希望它看起来如何的示例,它会更有帮助
  • 向宽格式添加行标识符,例如iris %&gt;% mutate(i = row_number()) %&gt;% gather(var, val, -i) %&gt;% spread(var, val)
  • 感谢@alistaire 的评论!

标签: r syntax-error spread


【解决方案1】:

稍后添加:抱歉@alistaire,在发布此回复后才看到您对原始帖子的评论。

据我了解Error: Duplicate identifiers for rows...,当您具有具有相同标识符的值时会发生这种情况。例如在原始 'iris' 数据集中,前五行 Species = setosa 的 Petal.Width 均为 0.2,三行 Petal.Length 有值1.4。收集这些数据不是问题,但是当您尝试传播它们时,该函数不知道什么属于什么。即0.2Petal.Width和1.4Petal.Length属于setosa的哪一行。

我在这些情况下使用的(tidyverse)解决方案是在收集阶段为每行数据创建一个唯一标记,以便该函数可以在您想要再次传播时跟踪哪些重复数据属于哪些行。请参见下面的示例:


# Load packages
library(dplyr)
library(tidyr)

# Get data
dataset <- iris

# View dataset
head(dataset)
#>   Sepal.Length Sepal.Width Petal.Length Petal.Width Species
#> 1          5.1         3.5          1.4         0.2  setosa
#> 2          4.9         3.0          1.4         0.2  setosa
#> 3          4.7         3.2          1.3         0.2  setosa
#> 4          4.6         3.1          1.5         0.2  setosa
#> 5          5.0         3.6          1.4         0.2  setosa
#> 6          5.4         3.9          1.7         0.4  setosa

# Gather data
dataset_gathered <- dataset %>% 
    # Create a unique identifier for each row 
    mutate(marker = row_number(Species)) %>%
    # Gather the data
    gather(key = Type, value = Values, 1:4)

# View gathered data
head(dataset_gathered)
#>   Species marker         Type Values
#> 1  setosa      1 Sepal.Length    5.1
#> 2  setosa      2 Sepal.Length    4.9
#> 3  setosa      3 Sepal.Length    4.7
#> 4  setosa      4 Sepal.Length    4.6
#> 5  setosa      5 Sepal.Length    5.0
#> 6  setosa      6 Sepal.Length    5.4

# Spread it out again
dataset_spread <- dataset_gathered %>%
    # Group the data by the marker
    group_by(marker) %>%
    # Spread it out again
    spread(key = Type, value = Values) %>%
    # Not essential, but remove marker
    ungroup() %>%
    select(-marker)

# View spread data
head(dataset_spread)
#> # A tibble: 6 x 5
#>   Species Petal.Length Petal.Width Sepal.Length Sepal.Width
#>    <fctr>        <dbl>       <dbl>        <dbl>       <dbl>
#> 1  setosa          1.4         0.2          5.1         3.5
#> 2  setosa          1.4         0.2          4.9         3.0
#> 3  setosa          1.3         0.2          4.7         3.2
#> 4  setosa          1.5         0.2          4.6         3.1
#> 5  setosa          1.4         0.2          5.0         3.6
#> 6  setosa          1.7         0.4          5.4         3.9

(一如既往,感谢 Jenny Bryan 提供了 reprex 软件包)

【讨论】:

  • 感谢彼得的精彩回答
【解决方案2】:

我们可以通过data.table 做到这一点

library(data.table)
dcast(melt(setDT(dataset, keep.rownames = TRUE), id.var = c("rn", "Species")), rn + Species ~ variable)

【讨论】:

  • 感谢您的回答
猜你喜欢
  • 1970-01-01
  • 2016-09-16
  • 2016-05-24
  • 1970-01-01
  • 2018-10-18
  • 1970-01-01
  • 1970-01-01
  • 2016-11-25
  • 1970-01-01
相关资源
最近更新 更多