【问题标题】:Add new column to a data frame based on an existing value根据现有值向数据框添加新列
【发布时间】:2017-07-11 16:31:06
【问题描述】:

我需要一种方法来根据 target_id 过滤我的数据。因为我有一组 1600 个没有一致名称的 target_id 值,而另一组包含单词“comp”,所以我认为创建一个具有基于 target_id 值的值的新列可能是最简单的。我有一个有一百万行的数据框,看起来像这样(只是抓取随机行来显示它的要点):

      sample_id          target_id l ength eff_length est_counts     tpm
159  SRR3884838C           CR1_Mam   2204       2005           0           0
160  SRR3884838C         CYRA11_MM    617        418           0           0
161  SRR3884838C          DERV2a_I   5989       5790          19    0.734541
162  SRR3884838C        DERV2a_LTR    335        136           7     11.5213
1094236 SRR3884878C comp78901_c0_seq3_1 1115     916       113.4     32.3604
1094237 SRR3884878C comp85230_c0_seq1_1 1201     1002      514       134.088
1094238 SRR3884878C comp56944_c0_seq1_1 2484     2285      10.5      1.20115

我需要创建一个新列(“类”),其中包含“comp”的 sample_ids 的值为 1,所有其他列的值为 0。这可能吗?数据有 40 个样本(SRR3884838 --> SRR3884878),每个样本都有相同的 target_ids 集,一组不统一的目标名称,然后是另一组都包含 comp。示例(由于格式原因删除了 tpm 列)

 sample_id          target_id       length   eff_length      est_counts class
159  SRR3884838C           CR1_Mam   2204       2005           0           0        
160  SRR3884838C         CYRA11_MM    617        418           0           0
161  SRR3884838C          DERV2a_I   5989       5790          19           0
162  SRR3884838C        DERV2a_LTR    335        136           7           0
1094236 SRR3884878C comp78901_c0_seq3_1 1115     916       113.4           1
1094237 SRR3884878C comp85230_c0_seq1_1 1201     1002      514             1
1094238 SRR3884878C comp56944_c0_seq1_1 2484     2285      10.5            1

我尝试使用合并函数,首先创建一个新的数据框,该数据框的类列具有一组 target_ids 的正确值,可能不正确的期望是它将创建新列,其中一个 target_ids 的实例已列出,但是当我这样做时,它删除了 eff_length 列并弄乱了数据的格式。我发现的所有示例中,用户根据使用数字的另一列值创建新列,但我不知道如何使用字符串 comp 来实现。这是我所做的:

total <- merge(data frameA,data frameB,by="target_id")

df A 是我的原始数据,而 df B 看起来像上面带有类列的示例。

【问题讨论】:

  • df$class &lt;- grepl('comp', df$taget_id) 会给出一个逻辑向量;将grepl 部分包裹在as.numeric 或as.integer 中,得到一个由0 和1 组成的向量。

标签: r


【解决方案1】:

使用:

df$class <- as.integer(grepl('comp', df$target_id))

给予:

> df
          sample_id           target_id length eff_length est_counts class
159     SRR3884838C             CR1_Mam   2204       2005        0.0     0
160     SRR3884838C           CYRA11_MM    617        418        0.0     0
161     SRR3884838C            DERV2a_I   5989       5790       19.0     0
162     SRR3884838C          DERV2a_LTR    335        136        7.0     0
1094236 SRR3884878C comp78901_c0_seq3_1   1115        916      113.4     1
1094237 SRR3884878C comp85230_c0_seq1_1   1201       1002      514.0     1
1094238 SRR3884878C comp56944_c0_seq1_1   2484       2285       10.5     1

【讨论】:

  • 如果我做对了,如果该 ID 的任何行中有"comp",OP 想要为样本 ID 获取 1。不确定。
  • 输出看起来像我想要的,但是当我复制并粘贴你写的内容时,我得到了错误:$&lt;-.data.frame(*tmp*, "class", value = integer(0 )) :替换有 0 行,数据有 1094241 我的数据框被称为 df 所以我不确定我哪里出错了。抱歉,我对此很陌生,也很愚蠢
  • @ZincFingers 你能提供一些重现问题的示例数据吗?
  • 我从一个导入的 .tsv 文件创建了我的数据框。您在答案中称为 df 的数据是示例数据。我将您答案中的表格复制并粘贴到一个新的 .tsv 文件中,并像这样创建了一个 testdf:testdf &lt;- read.table("filelocation/Test.tsv")testdf$class &lt;- as.integer(grepl('comp', df$target_id))
  • 哦。出于某种原因,我认为 header = TRUE 是默认值。您的修复有效,我现在已正确格式化数据。感谢您的帮助。我也会确保查看您提供的链接。
【解决方案2】:

sample$class &lt;- as.numeric(grepl ("^comp", sample$target_id)) 怎么样?

【讨论】:

  • 已编辑。我认为鉴于 TRUE/FALSE 和 1/0 之间的可互换性,没有必要更改为数字。
  • sample$target_id 中的错误:'closure' 类型的对象不是子集
  • 我尝试粘贴您的 df,但我的代码没有发现任何问题。 > test$class test sample_id target_id length eff_length est_counts class 1 SRR3884838C CR1_Mam 2204 2005 0.0 0
猜你喜欢
  • 2015-11-19
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-12-29
  • 1970-01-01
  • 2020-02-24
  • 1970-01-01
  • 2019-09-29
相关资源
最近更新 更多