【问题标题】:I want to transform an R dataframe into a dataframe with a column of list [duplicate]我想将 R 数据框转换为具有列表列的数据框 [重复]
【发布时间】:2020-07-14 01:51:28
【问题描述】:

好的,我有一个两列 data.frame,可变编号为 child 到 head。 (其他2列是参考)

dat <- read.table(header=TRUE, stringsAsFactors=FALSE, text="
                        head          child UID     logic
                    1  01001         01001   1     FALSE
                    2  01001         01021   2      TRUE
                    3  01001         01047   3      TRUE
                    4  01001         01051   4      TRUE
                    5  01001         01085   5      TRUE
                    6  01001         01101   6      TRUE
                    7  01003         01003   7     FALSE
                    8  01003         01025   8      TRUE
                    9  01003         01053   9      TRUE
                    10 01003         01097  10      TRUE
                    11 01003         01099  11      TRUE
                    12 01003         01129  12      TRUE
                    13 01003         12033  13      TRUE
                    14 01005         01005  14     FALSE
                    15 01005         01011  15      TRUE
                    16 01005         01045  16      TRUE
                    17 01005         01067  17      TRUE
                    18 01005         01109  18      TRUE
                    19 01005         01113  19      TRUE
                    20 01005         13061  20      TRUE
                    21 01005         13239  21      TRUE
                    22 01005         13259  22      TRUE")

我希望唯一的 head 只有三行,child 的列表。

如果您有更好的方法建议,我愿意接受。

其他列UID和logic我已经添加以供参考,但可以删除它们。

在我的尝试中,我尝试转换为带有边缘列表的图形,然后转换为 JSON。

# make graph ##########
library(tidyverse)
library(igraph)
library(jsonlite)
gdat <- select(dat, head, child)
mdat <- as.matrix(gdat)
edge_dat <- graph_from_edgelist(mdat)
plot.igraph(edge_dat)
jdat <- toJSON(mdat, matrix = "rowmajor")

期望的输出:

head   child1   child2   child3   child4   child5   child6   child7
01001  01001    01021    01047    01051    01085    01101    NA
01003  01003    01025    01053    ... and so on
01005  01005    01011    ... and so on

【问题讨论】:

    标签: r data.table tidyr adjacency-matrix jsonlite


    【解决方案1】:

    这是你想要的吗?

    setDT(dat)
    
    dat_child <- dat[(logic)]
    dat_child[,.(list(unique(child))), by = "head"]
    
    dat_child
       head                                      V1
    1: 1001                1021,1047,1051,1085,1101
    2: 1003      1025, 1053, 1097, 1099, 1129,12033
    3: 1005  1011, 1045, 1067, 1109, 1113,13061,...
    

    【讨论】:

    • 是的,我想是的...您是否推荐任何其他最好的方式来存储这样的数据?
    • 您有很多选择,具体取决于您以后要如何处理它们。您可以使用建议的列表列、标准列表(改用split(dat_child, by = "head"))、嵌套数据(嵌套data.table 或嵌套tibble)。这确实是您稍后要做的,它将确定最合适的结构
    猜你喜欢
    • 2020-05-31
    • 1970-01-01
    • 1970-01-01
    • 2016-07-24
    • 1970-01-01
    • 2015-11-14
    • 1970-01-01
    • 2015-04-16
    • 1970-01-01
    相关资源
    最近更新 更多