【发布时间】:2015-02-16 02:41:00
【问题描述】:
我有一个用例,我需要拆分 data.table,然后对每个分区应用不同的按引用修改操作。但是,拆分会强制复制每个表。
这是一个关于 iris 数据集的玩具示例:
#split the data
DT <- data.table(iris)
out <- split(DT, DT$Species)
#assign partitions to global environment
NAMES <- as.character(unique(DT$Species))
lapply(seq_along(out), function(x) {
assign(NAMES[x], out[[x]], envir=.GlobalEnv)})
#modify by reference, same function applied to different columns for different partitions
#would do this programatically in real use case
virginica[ ,summ:=sum(Petal.Length)]
setosa[ ,summ:=sum(Petal.Width)]
#rbind all (again, programmatic)
do.call(rbind, list(virginica, setosa))
然后我收到以下警告:
Warning message:
In `[.data.table`(out$virginica, , `:=`(cumPedal, cumsum(Petal.Width))) :
Invalid .internal.selfref detected and fixed by taking a copy of the whole table so that := can add this new column by reference.
我知道这与将 data.tables 放入列表有关。这个用例有什么解决方法,或者有办法避免使用split?请注意,在实际情况下,我想以编程方式通过引用进行修改,因此对解决方案进行硬编码是行不通的。
【问题讨论】:
-
您不需要
split和data.table。您可能正在寻找data.table中的.EACHI函数。 -
问题是
j表达式对于每个分区实际上是不同的。我正在以编程方式为每个分区构建不同的j表达式。 -
我将更新示例以显示我的目标。
-
所以,明确地说,您正在寻找每个组来创建一个新的(不同名称的)列,并将函数应用于不同的列?还是功能也完全不同?
-
那么我认为
.EACHI会起作用。给我几分钟从手机切换到电脑:-)
标签: r data.table