【问题标题】:Apply a function on a data-frame and return a data-frame对数据框应用函数并返回数据框
【发布时间】:2018-06-21 08:40:31
【问题描述】:

我有一个这样的数据框

 ID  07  08  09  10  year balance

abc   0   0   0   0  09    2123.00
efg   0   0   0   0  09    780.4
xyz   0   0   0   0  07    2402.9
prq   0   0   0   0  10    123.3
mno   0   0   0   0  07    679

我需要根据“year”列和balance中的值填写07、08、09和10列。 对于每个 ID,对应列年份中的值的列填充余额中的值。逐行应用。

例如,对于第 1 行,年份是 09,因此该 ID 的第 09 列用值 2123.00 填充。其余年份值保持为 0。

对于第 3 行,值 24502.9 填充到第 07 列,因为它的年份值是 07。以此类推..

我的输出应该是这样的

 ID  07      08  09      10    year  balance

abc   0      0  2123.00  0      09    2123.00
efg   0      0  780.4    0      09    780.4
xyz  2402.9  0   0       0      07    2402.9
prq   0      0   0      123.3   10    123.3
mno  679     0   0       0      07    679

PS:我已经为此编写了一个 for 循环。我需要比 for 循环更快的东西。我实际上正在处理成千上万的数据。 我不知道是否有类似 apply 的东西返回一个数据框

【问题讨论】:

  • 也许输入你的数据框?
  • 您可以使用 ID、年份、data.frame 的余额列和 dcast 它使用 reshape2 和 data.table 以 ID 作为列中的行和年份,并在值中使用余额library(reshape2) library(data.table) final_output<-dcast(setDT(df),ID~year, value.vars="balance")
  • @SatZ 这似乎确实有效。请把它写成答案
  • 如果我投的是年份,年份会按升序排列吗?我需要这样
  • @DomJo 我已将其添加为答案以及重新排序

标签: r dataframe


【解决方案1】:

基本上你想要做的是将数据框的右侧从长格式转换为宽格式。您可以使用tidyr 中的spread 函数来执行此操作。

library(tidyr)
library(dplyr)

D <- read.table(header=TRUE, text="
ID  07  08  09  10  year balance
abc  0   0   0   0  09    2123.00
efg  0   0   0   0  09    780.4
xyz  0   0   0   0  07    24502.9
prq  0   0   0   0  10    123.3
mno  0   0   0   0  07    679")

D %>%
  select(ID, year, balance) %>%
  spread(year, balance, fill=0) %>%
  bind_cols(D[,c("year","balance")])

#>    ID       7      9    10 year balance
#> 1 abc     0.0 2123.0   0.0    9  2123.0
#> 2 efg     0.0  780.4   0.0    9   780.4
#> 3 mno   679.0    0.0   0.0    7 24502.9
#> 4 prq     0.0    0.0 123.3   10   123.3
#> 5 xyz 24502.9    0.0   0.0    7   679.0

注意:输出中缺少 08 年,因为您的示例数据中缺少它。

【讨论】:

  • 这也将折叠行,这样ID 只能得到一行,并且可能会连续几年,这很可能是可取的(但随后bind_cols 会给出意想不到的结果) .要保留您可以执行的行数:df1[c("ID","year","balance")] %&gt;% mutate(year2=year,balance2=balance) %&gt;% spread(year2,balance2)
  • 我只显示了前几条记录。 08 的数据如下所示。 08、07 等表示 2008、2007 等。如果有新数据进入,年份列可能会发生变化。所以我不能按名称引用
【解决方案2】:

我确定你想要这个

do.call(rbind, lapply(1:nrow(df1), function(i) {
  df1[i, df1[i, 6]] <- df1[i, 7] 
  df1[i, ]
  }))

产量

   ID     07 08     09    10 year balance
1 abc    0.0  0 2123.0   0.0   09  2123.0
2 efg    0.0  0  780.4   0.0   09   780.4
3 xyz 2402.9  0    0.0   0.0   07  2402.9
4 prq    0.0  0    0.0 123.3   10   123.3
5 mno  679.0  0    0.0   0.0   07   679.0

数据

df1 <- structure(list(ID = structure(c(1L, 2L, 5L, 4L, 3L), .Label = c("abc", 
"efg", "mno", "prq", "xyz"), class = "factor"), `07` = c(0L, 
0L, 0L, 0L, 0L), `08` = c(0L, 0L, 0L, 0L, 0L), `09` = c(0L, 0L, 
0L, 0L, 0L), `10` = c(0L, 0L, 0L, 0L, 0L), year = c("09", "09", 
"07", "10", "07"), balance = c(2123, 780.4, 2402.9, 123.3, 679
)), row.names = c(NA, -5L), class = "data.frame")

【讨论】:

    【解决方案3】:

    您可以使用data.table 和reshape2 包来执行此操作。

    您可以使用 data.frame 的 ID、年份、余额列和 dcast,其中 ID 作为列中的行和年份,并在值中使用余额

    library(reshape2) 
    library(data.table) 
    final_output<-dcast(setDT(df),ID~year, value.var="balance")
    

    如果您想对列重新排序,可以使用以下参考中的 sn-p: Reordering dcast data frame

    final_output<-dcast(setDT(df),ID~reorder(year,year), value.var="balance")
    

    【讨论】:

      【解决方案4】:

      你可以使用 4 行:

      df$`07` <- ifelse(test = df$year=='07',yes = df$balance, no=0)
      df$`08` <- ifelse(test = df$year=='08',yes = df$balance, no=0) 
      df$`09` <- ifelse(test = df$year=='09',yes = df$balance, no=0)
      df$`10` <- ifelse(test = df$year=='10',yes = df$balance, no=0)
      

      我认为与循环相比,它的工作速度会超快

      【讨论】:

      • 老兄。 07, 08 .. 可能因数据而异。这需要灵活。所以我不喜欢按名称引用
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2016-08-27
      • 1970-01-01
      • 1970-01-01
      • 2018-04-26
      • 1970-01-01
      • 1970-01-01
      • 2021-11-23
      相关资源
      最近更新 更多