【问题标题】:R Grouped Data Frame: Function relates single value to the other values of the groupR分组数据框:函数将单个值与组的其他值相关联
【发布时间】:2018-12-02 15:38:57
【问题描述】:

在分组数据框中,我想应用一个函数,该函数将实际 a 行中的一个值与该组(和同一列)的所有其他值(当前行中的一个 i 除外)相关联。这将导致一个单值新变量。因此,如果该组由 c(1,2,3,4,5) 组成,我希望有一个新变量: c(fun(1,c(2,3), fun(2, c(1,3) ), 乐趣(3, c(1,2)) 我的小组没有类似的规模。但是尝试了这么久,我总是收到一些有趣的值,比如零或错误。

示例代码:

  set.seed(3)  
dat <- data_frame(a=1:10,value=round(runif(10),2),group=c(1,1,1,2,2,3,3,3,3,4))

 # one possible function
dif.dist <- function(x1, x2) sum(abs(x1 - x2))/(length(x2)-1) 

 # with this, sometimes the grouping gets lost in "vec" and i receive zeros   
 x <- dat%>%
 group_by(group)%>%
 mutate(vec= list(value))%>%
 mutate(dif = dif.dist(unique(value),unlist(vec)[unlist(vec)!=value]))%>%
 ungroup()

 # another try with plyr, that returns only 0   
 dat <- ddply(dat, .(group), mutate, dif=dif.dist1(value[a==a],value[value!=value[a==a]]))

但功能有效

  dif.dist(dat$value[1],dat$value[2:3])
 [1] 0.85

稍后,我需要它来接收与每个参与者相关的大量变量的距离矩阵。非常感谢您的帮助!

【问题讨论】:

    标签: r dplyr grouped-table


    【解决方案1】:

    一种选择是在按“组”分组后循环遍历行序列,并根据索引对“值”的元素进行子集化

    library(dplyr)
    library(purrr)
    out <- dat %>%
             group_by(group) %>% 
             mutate(dif = map_dbl(row_number(), ~ dif.dist(value[.x], value[-.x])))
    
    head(out, 2)
    # A tibble: 2 x 4
    # Groups:   group [1]
    #      a value group   dif
    #  <int> <dbl> <dbl> <dbl>
    #1     1  0.17     1  0.85
    #2     2  0.81     1  1.07
    

    【讨论】:

    • 非常感谢,我不知道 purrr。虽然:当我尝试它时,一条错误消息显示:“Fehler in rank(x, ties.method = "first", na.last = "keep"):缺少参数 "x" (ohne Standardwert)"。您是否在数据中添加了某物?还是我需要为x填写某事?谢谢你的提示!
    • @SansSoleil 我只使用了你的数据。没有收到任何错误。可能是因为包的版本不同?
    • 是的,需要更新,现在一切运行良好-再次感谢您!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2012-02-05
    • 2017-10-20
    • 1970-01-01
    • 2019-04-02
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多