【发布时间】:2015-10-11 14:44:51
【问题描述】:
我有一个格式如下的数据框:
id | name | logs
---+--------------------+-----------------------------------------
84 | "zibaroo" | "C47931038"
12 | "fabien kelyarsky" | c("C47331040", "B19412225", "B18511449")
96 | "mitra lutsko" | c("F19712226", "A18311450")
34 | "PaulSandoz" | "A47431044"
65 | "BeamVision" | "D47531045"
如您所见,“日志”列包含每个单元格中的字符串向量。
是否有一种有效的方法可以将数据帧转换为长格式(每行一个观察值),而无需将“日志”分成几列的中间步骤?
这很重要,因为数据集非常大,而且每个人的日志数量似乎是任意的。
换句话说,我需要以下内容:
id | name | log
---+--------------------+------------
84 | "zibaroo" | "C47931038"
12 | "fabien kelyarsky" | "C47331040"
12 | "fabien kelyarsky" | "B19412225"
12 | "fabien kelyarsky" | "B18511449"
96 | "mitra lutsko" | "F19712226"
96 | "mitra lutsko" | "A18311450"
34 | "PaulSandoz" | "A47431044"
65 | "BeamVision" | "D47531045"
这是真实数据框的一部分的dput:
structure(list(id = 148:157, name = c("avihil1", "Niarfe", "doug henderson",
"nick tan", "madisp", "woodbusy", "kevinhcross", "cylol", "andrewarrow",
"gstavrev"), logs = list("Z47331572", "Z47031573", c("F47531574",
"B195945", "D186871", "S192939", "S182865", "G19539045"), c("A47231575",
"A190933", "C181859"), "F47431576", c("B47231577", "D193936",
"Q184862"), "Y47331579", c("A47531580", "Z195944", "B185870"),
"N47731581", "E47231582")), .Names = c("id", "name", "logs"
), row.names = 149:158, class = "data.frame")
【问题讨论】:
-
请提供您的(一段)数据的
dput。 -
这真的是文本文件的样子,还是您对 R-data-object 可能包含的内容的想象?如果是后者,那么您应该发布
dput(object)的输出。 -
你可能需要
library(splitstackshape); cSplit(data, 'log', ',', direction = 'long') -
@SabDeM,@BondedDuest:我添加了
dput