【问题标题】:splitting a data.table, then modifying by reference拆分data.table,然后通过引用修改
【发布时间】:2015-02-16 02:41:00
【问题描述】:

我有一个用例,我需要拆分 data.table,然后对每个分区应用不同的按引用修改操作。但是,拆分会强制复制每个表。

这是一个关于 iris 数据集的玩具示例:

#split the data
DT <- data.table(iris)
out <- split(DT, DT$Species)

#assign partitions to global environment
NAMES <- as.character(unique(DT$Species))
lapply(seq_along(out), function(x) {
assign(NAMES[x], out[[x]], envir=.GlobalEnv)})

#modify by reference, same function applied to different columns for different partitions
#would do this programatically in real use case
virginica[ ,summ:=sum(Petal.Length)]
setosa[ ,summ:=sum(Petal.Width)]

#rbind all (again, programmatic)
do.call(rbind, list(virginica, setosa))

然后我收到以下警告:

 Warning message:
 In `[.data.table`(out$virginica, , `:=`(cumPedal, cumsum(Petal.Width))) :
  Invalid .internal.selfref detected and fixed by taking a copy of the whole table so that := can add this new column by reference.

我知道这与将 data.tables 放入列表有关。这个用例有什么解决方法,或者有办法避免使用split?请注意,在实际情况下,我想以编程方式通过引用进行修改,因此对解决方案进行硬编码是行不通的。

【问题讨论】:

  • 您不需要split 和data.table。您可能正在寻找data.table 中的.EACHI 函数。
  • 问题是j表达式对于每个分区实际上是不同的。我正在以编程方式为每个分区构建不同的 j 表达式。
  • 我将更新示例以显示我的目标。
  • 所以,明确地说,您正在寻找每个组来创建一个新的(不同名称的)列,并将函数应用于不同的列?还是功能也完全不同?
  • 那么我认为.EACHI 会起作用。给我几分钟从手机切换到电脑:-)

标签: r data.table


【解决方案1】:

下面是一个使用.EACHI 来实现听起来像你想要做的事情的例子:

## Create a data.table that indicates the pairs of keys to columns
New <- data.table(
  Species = c("virginica", "setosa", "versicolor"), 
  FunCol = c("Petal.Length", "Petal.Width", "Sepal.Length"))

## Set the key of your original data.table
setkey(DT, Species)

## Now use .EACHI
DT[New, temp := cumsum(get(FunCol)), by = .EACHI][]
#      Sepal.Length Sepal.Width Petal.Length Petal.Width   Species  temp
#   1:          5.1         3.5          1.4         0.2    setosa   0.2
#   2:          4.9         3.0          1.4         0.2    setosa   0.4
#   3:          4.7         3.2          1.3         0.2    setosa   0.6
#   4:          4.6         3.1          1.5         0.2    setosa   0.8
#   5:          5.0         3.6          1.4         0.2    setosa   1.0
#  ---                                                                  
# 146:          6.7         3.0          5.2         2.3 virginica 256.9
# 147:          6.3         2.5          5.0         1.9 virginica 261.9
# 148:          6.5         3.0          5.2         2.0 virginica 267.1
# 149:          6.2         3.4          5.4         2.3 virginica 272.5
# 150:          5.9         3.0          5.1         1.8 virginica 277.6

## Basic verification
head(cumsum(DT["setosa", ]$Petal.Width), 5)
# [1] 0.2 0.4 0.6 0.8 1.0
tail(cumsum(DT["virginica", ]$Petal.Length), 5)

【讨论】:

  • 这个解决方案可以工作,但是被调用的函数必须是矢量化的,因为我们不能设置 = 唯一键
猜你喜欢
  • 2019-12-06
  • 1970-01-01
  • 2020-12-17
  • 2021-05-03
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2012-06-04
  • 1970-01-01
相关资源
最近更新 更多