【问题标题】:R: extracting pattern, different timesR:提取模式,不同时间
【发布时间】:2017-04-12 08:19:17
【问题描述】:

我有以下问题:我有一个文本,由章节分隔并由向量存储。假设是这样的:

text <- c("Here are information about topic1.", 
"Here are some information about topic2 or topic3.", 
"Chapter number 4 is really annoying.", 
"Topic4 is discussed in this chapter.")

我想提取不同章节中提到的不同主题。所以我的输出应该是这样的:

output
      [1]       [2]
[1] "topic1"
[2] "topic2" "topic3"
[3]
[4] "topic3"

所以我有一些行有多个结果,而有些行没有匹配。

我尝试使用 str_extract_all 并取消列出列表,但遇到了导致行元素数量不同的问题。

谢谢大家!

【问题讨论】:

    标签: r string extract


    【解决方案1】:

    您可以从plyr 使用rbind.fill.matrix。

    text <- c("Here are information about topic1.", 
              "Here are some information about topic2 or topic3.", 
              "Chapter number 4 is really annoying.", 
              "Topic4 is discussed in this chapter.")
    
    library(stringr)
    library(plyr)
    
    xy <- str_extract_all(text, pattern = "[Tt]opic\\d+")
    xy <- sapply(xy, FUN = function(x) matrix(x, nrow = 1))
    rbind.fill.matrix(xy) # from plyr
    
         1        2       
    [1,] "topic1" NA      
    [2,] "topic2" "topic3"
    [3,] NA       NA      
    [4,] "Topic4" NA 
    

    【讨论】:

      猜你喜欢
      • 2020-11-02
      • 1970-01-01
      • 2018-08-16
      • 2021-06-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-05-20
      • 2014-02-12
      相关资源
      最近更新 更多