【问题标题】:Eliminating more than one succeeding instance of a string down cells in R在 R 中消除多个连续的字符串下单元实例
【发布时间】:2020-08-06 00:27:08
【问题描述】:

我对 R 比较陌生。我有一个包含 500 万个观察值和 1 个变量的数据框,看起来像这样:

PMID- 28524368 PMID- 28504342 PMID- 28501042 RN - 4964P6T9RB (Aldosterone) RN - EC 3.4.23.15 (Renin) RN - RWP5GA015D (Potassium) MH - Adrenal Cortex Neoplasms/*diagnostic imaging/pathology/surgery MH - Adrenocortical Adenoma/*diagnostic imaging/pathology/surgery MH - Aldosterone/blood MH - Humans PMID- 28523858 PMID- 28517030 PMID- 28513869 MH - Hyperaldosteronism/*complications MH - Hypertension/*etiology MH - Male MH - Middle Aged MH - Potassium/blood PMID- 28494487 PMID- 28493475 MH - Renin/blood MH - Tomography, X-Ray Computed

但是,我只希望连续有 1 个 PMID,第一个 - 其余的 PMID 应该被删除,导致数据帧看起来像:

PMID- 28524368 RN - 4964P6T9RB (Aldosterone) RN - EC 3.4.23.15 (Renin) RN - RWP5GA015D (Potassium) MH - Adrenal Cortex Neoplasms/*diagnostic imaging/pathology/surgery MH - Adrenocortical Adenoma/*diagnostic imaging/pathology/surgery MH - Aldosterone/blood MH - Humans PMID- 28523858 MH - Hyperaldosteronism/*complications MH - Hypertension/*etiology MH - Male MH - Middle Aged MH - Potassium/blood PMID- 28494487 MH - Renin/blood MH - Tomography, X-Ray Computed

请指教。我尝试使用:

# remove excessive PMIDs
for (i in nrow(original_reduced))
{
  if (substr(original_reduced[i, 1], 1, 4) == "PMID")
  {
    if (substr(original_reduced[i+1, 1], 1, 4) == "PMID" && i != nrow(original_reduced)) # if next row is also PMID
    {
      original_reduced <- original_reduced[-c(i+1), ] # delete entry after
    }
  }
}

但我收到了这个错误:

Error in if (substr(original_reduced[i + 1, 1], 1, 4) == "PMID") { : missing value where TRUE/FALSE needed

即使我的数据框中没有 NA。

谢谢。

【问题讨论】:

    标签: r string dataframe substr


    【解决方案1】:

    这是一个可行的解决方案。代码解释见 cmets

     df<-structure(list(V1 = c("PMID- 28524368", "PMID- 28504342", "PMID- 28501042", 
    "RN - 4964P6T9RB (Aldosterone)", "RN - EC 3.4.23.15 (Renin)", 
    "RN - RWP5GA015D (Potassium)", "MH - Adrenal Cortex Neoplasms/*diagnostic imaging/pathology/surgery", 
    "MH - Adrenocortical Adenoma/*diagnostic imaging/pathology/surgery", 
    "MH - Aldosterone/blood", "MH - Humans", "PMID- 28523858", "PMID- 28517030", 
    "PMID- 28513869", "MH - Hyperaldosteronism/*complications", "MH - Hypertension/*etiology", 
    "MH - Male", "MH - Middle Aged", "MH - Potassium/blood", "PMID- 28494487", 
    "PMID- 28493475", "MH - Renin/blood", "MH - Tomography, X-Ray Computed"
    )), .Names = "V1", row.names = c(NA, -22L), class = "data.frame")
    
    library(dplyr)
    
    #Add flag for PMID rows
       df$pmid<-grepl("^PMID", df$V1)
    #find rows of where n == n+1
       matches<-df$pmid==lag(df$pmid)
    #find rows equal to previous row and is a PMID row
       toremove<-which(matches==TRUE & df$pmid==TRUE)
    #remove rows
       df<-df[-toremove,]
       df$pmid<-NULL  #remove added column
    

    【讨论】:

    • 您的代码删除了所有的 PMID...我假设 lag 函数没有正常工作...看起来很有希望,但是
    • 是否加载了 dplyr 库?如果使用基础包的 lag 函数,这将不起作用。
    • 我通过将 n = 1 添加到您的 lag 函数来修复它...谢谢! matches &lt;-df$pmid == lag(df$pmid, n = 1)。也非常感谢您的 cmets - 作为初学者,他们真的帮助了我!
    【解决方案2】:

    试试这个:

    df%>%mutate(number=sequence(rle(name)[['lengths']]))%>%filter((number==1 & grepl('PMID',number))|!grepl('PMID',name))%>%select(name)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2019-01-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2016-07-31
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多