【问题标题】:How to identify repeated subsequences in a dataset如何识别数据集中的重复子序列
【发布时间】:2019-01-19 12:06:35
【问题描述】:

我有一个数值数据集,每个数值代表一个区域。

例如。

x <- c(1,6,1,2,3,4,5,8,5,9,10,1,2,3,10,7,5,9,4,1,2,3)

我需要确定数据中是否存在重复的子序列,即受试者是否重复从区域 1 到区域 2 到 3。在上面的示例中,1,2,3 将给出值 3。我没有已经知道子序列,我需要 R 提供给定数据。

接下来我需要计算这个子序列在数据中出现的次数。

非常基础的知识或 R 如果这是一个简单的任务,请原谅我的无知!

【问题讨论】:

  • 会有这样的工作吗? library(stringr);table(gsub("_","",unlist(str_extract_all(str_c(x,collapse = "_"),"(\\w{4,})(?=.*\\1)")))) + 1???

标签: r subsequence


【解决方案1】:

这是一种查找长度为 n 的序列以及重复次数的方法

对于n = 3

library(tidyverse) # not necessary, see base version below

n <- 3
lapply(seq(0, length(x) - n), `+`, seq(n)) %>% # get index of all subsequences
  map_chr(~ paste(x[.], collapse = ',')) %>% # paste together as character
  table %>% # get number of times each occurs
  `[`(. > 1) # select sequences occurring > 1 time
# 1,2,3 
# 3 

对于n = 2

n <- 2
lapply(seq(0, length(x) - n), `+`, seq(n)) %>% 
  map_chr(~ paste(x[.], collapse = ',')) %>% 
  table %>% 
  `[`(. > 1)
# 1,2 2,3 5,9 
# 3   3   2 

没有 Tidyverse

seqs <- lapply(seq(0, length(x) - n), `+`, seq(n))
seqs.char <- sapply(seqs, function(i) paste(x[i], collapse = ','))
tbl <- table(seqs.char)
tbl[tbl > 1]

我将添加我自己的问题:有没有人知道如何在不先转换为字符的情况下做到这一点?例如fun 其中fun(list(1:2, 1:2, 2:3)) 告诉您1:2 出现两次而2:3 出现一次?

【讨论】:

  • 对不起,我对 R 真的很陌生!我正在使用 tidyverse。但我收到以下错误: lapply(c("layla", seq(length(x) - n)), +, seq(n)) %>% map_chr(~paste(x[.], : 找不到函数“%>%”
  • @Melanie %&gt;% 函数(称为“管道”)在默认情况下不是 R 的一部分,但与 tidyverse 库(等等)一起加载。该错误告诉您它找不到%&gt;%,因为未加载tidyverse。您可以运行install.packages('tidyverse'),然后运行library(tidyverse) 以加载tidyverse,然后再运行代码。另一种选择是使用我的其他方法,它给出相同的结果并且不需要tidyverse。 (在代码块中显示“没有 Tidyverse”)
【解决方案2】:

另一种tidyverse 方法可根据您希望子序列具有多少值来创建结果的大数据框:

library(tidyverse)

# example vector
x <- c(1,6,1,2,3,4,5,8,5,9,10,1,2,3,10,7,5,9,4,1,2,3)

# function that gets as input number of consequtive elements in a subsequence
# and returns an ordered dataframe by counts of occurence
f = function(n) {

  data.frame(value = x) %>%               # get the vector x
    slice(1:(nrow(.)-n+1)) %>%            # remove values not needed from the end
    mutate(position = row_number()) %>%   # add position of each value
    rowwise() %>%                         # for each value/row
    mutate(vec = paste0(x[position:(position+n-1)], collapse = ",")) %>% # create subsequences as a string
    ungroup() %>%                         # forget the grouping
    count(vec, sort = T) }                # order by counts descending


2:5 %>%                    # specify how many values in your subsequences you want to investigate (let's say from 2 to 5)
  map_df(~ data.frame(NumElements = ., f(.))) %>%  # apply your function and keep the number values
  arrange(desc(n)) %>%     # order by counts descending
  tbl_df()                 # (only for visualisation purposes)


# # A tibble: 88 x 3
#   NumElements vec       n
#         <dbl> <chr> <int>
# 1           2 1,2       3
# 2           2 2,3       3
# 3           3 1,2,3     3
# 4           2 5,9       2
# 5           2 1,6       1
# 6           2 10,1      1
# 7           2 10,7      1
# 8           2 3,10      1
# 9           2 3,4       1
# 10          2 4,1       1
# # ... with 78 more rows

【讨论】:

    【解决方案3】:

    下面的方法找到任意长度的序列(k):输入向量被转换成一个有k行的矩阵;这是通过在开头添加0:(k-1) NA's 来完成k 次的。最后,计算这些k 矩阵中的所有行(paste'ing 元素在一起):

    frs <- function(x, k=2){
       padit <- function(.) c(.,rep(NA, k-length(.)%%k))
       xx <- lapply(1:k, function(iii) padit(c(rep(NA,iii-1), x)))
       xx <- do.call(rbind, lapply(xx, function(.) matrix(., ncol=k, byrow=TRUE)))
       xx <- sapply(split(xx, 1:NROW(xx)), paste, collapse=",")
       (function(x) x[x>1])(table(xx))
    

    }

    输出:

    > frs(x,2)
    xx
    1,2 2,3 5,9 
      3   3   2 
    > frs(x,3)
    1,2,3 
        3 
    > frs(x,4)
    named integer(0)
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2015-09-28
      • 1970-01-01
      • 2019-06-22
      • 2014-06-14
      • 1970-01-01
      • 1970-01-01
      • 2019-06-19
      • 1970-01-01
      相关资源
      最近更新 更多