【问题标题】:How to apply function to each row of data.table如何将函数应用于 data.table 的每一行
【发布时间】:2023-03-06 10:58:02
【问题描述】:

我正在尝试使用library(financial) 计算来自data.table 格式的给定现金流的每个观察值的净现值 (NPV)。这是我的现金流:

library(data.table)    
dt <- data.table(id=c(1,2,3,4), Year1=c(NA, 30, 40, NA), Year2=c(20, 30, 20 ,70), Year3=c(60, 40, 0, 10))

计算 NPV 并在 data.table 中更新,

library(financial)
npv <- apply(dt, 1, function(x) cf(na.omit(x[-1]), i = 20)$tab[, 'NPV'])
dt[, NPV:=npv]

返回,

   id Year1 Year2 Year3      NPV
1:  1    NA    20    60 70.00000
2:  2    30    30    40 82.77778
3:  3    40    20     0 56.66667
4:  4    NA    70    10 78.33333

如何使用函数 cf 将结果直接更新到 data.table 中的每一行?

仅供参考:在我的真实数据集中,有超过 50 列

【问题讨论】:

  • 也许dt[, "NPV" := apply(dt, 1, function(x) cf(na.omit(x[-1]), i = 20)$tab[, 'NPV'])] ??这是你想要的吗?
  • 或者dt[, NPV := cf(na.omit(c(Year1, Year2, Year3)), i=20)$tab[, 'NPV'], by=1:nrow(dt)]`

标签: r data.table


【解决方案1】:

我们可以尝试基于连接的方法

dt[melt(dt, id.var = "id")[, .(NPV = cf(value[!is.na(value)], 
                      i = 20)$tab[, "NPV"]), id], on = 'id']
#   id Year1 Year2 Year3      NPV
#1:  1    NA    20    60 70.00000
#2:  2    30    30    40 82.77778
#3:  3    40    20     0 56.66667
#4:  4    NA    70    10 78.33333

【讨论】:

  • 令人惊讶的是,它并没有带来那么大的速度优势。在这种情况下,cf 函数确实需要矢量化方法,否则无论如何它总是会运行nrow(dt) 次。
  • @thelatemail 在这种情况下,带有set 的for 循环可以提高速度。
【解决方案2】:

重写cf 函数以仅计算所需的部分将加快速度显着:

dt[, NPV := {x <- na.omit(unlist(.SD)); sum(x * sppv(20,0:(length(x)-1)))}, by=id]

#   id Year1 Year2 Year3      NPV
#1:  1    NA    20    60 70.00000
#2:  2    30    30    40 82.77778
#3:  3    40    20     0 56.66667
#4:  4    NA    70    10 78.33333

事实上,这可能现在可以向量化了……嗯,让我想想!

【讨论】:

    【解决方案3】:

    我们可以尝试制作自己的 npv 函数以用于此示例。

    dcf <- function(x, r, t0=FALSE){
      # calculates discounted cash flows (DCF) given cash flow and discount rate
      #
      # x - cash flows vector
      # r - vector or discount rates, in decimals. Single values will be recycled
      # t0 - cash flow starts in year 0, default is FALSE, i.e. discount rate in first period is zero.
      if(length(r)==1){
        r <- rep(r, length(x))
        if(t0==TRUE){r[1]<-0}
      }
      x/cumprod(1+r)
    }
    
    npv <- function(x, r, t0=FALSE){
      # calculates net present value (NPV) given cash flow and discount rate
      #
      # x - cash flows vector
      # r - discount rate, in decimals
      # t0 - cash flow starts in year 0, default is FALSE
      sum(dcf(x, r, t0))
    }
    

    现在,只要你想apply(x,1,f)、melt/gather/nest 代替。 除非您打算完全误导数据的用户,否则在计算 NPV 时,您永远不应放弃NA。这将意味着您将现金流折现到不同的时间点。将 NA 替换为 0。我还看到您打算使用的包将现金流量折现到第 0 年,基本上意味着第一个现金流量(在第 1 年)没有折现。

    library(data.table)
    npv_dt <- melt(dt, id.vars = "id")[is.na(value), value:=0][order(variable), .(NPV=npv(x=value, r=0.2, t0=TRUE)), by="id"]
    
    setkey(dt, id)
    setkey(npv_dt, id)
    
    npv_dt[dt]
    
    #>    id      NPV Year1 Year2 Year3
    #> 1:  1 58.33333    NA    20    60
    #> 2:  2 82.77778    30    30    40
    #> 3:  3 56.66667    40    20     0
    #> 4:  4 65.27778    NA    70    10
    

    【讨论】:

      猜你喜欢
      • 2013-03-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2015-12-15
      • 2017-07-01
      相关资源
      最近更新 更多