【问题标题】:Simple section labeling with tidytext for plain text input带有 tidytext 的简单部分标签,用于纯文本输入
【发布时间】:2017-02-23 21:11:41
【问题描述】:

我正在使用tidytext(和tidyverse)来分析一些文本数据(如Tidy Text Mining with R)。

我的输入文本文件 myfile.txt 如下所示:

# Section 1 Name
Lorem ipsum dolor
sit amet ... (et cetera)
# Section 2 Name
<multiple lines here again>

大约有 60 个部分。

我想生成一个列section_name,其中字符串"Category 1 Name" 或"Category 2 Name" 作为相应行的值。例如,我有

library(tidyverse)
library(tidytext)
library(stringr)

fname <- "myfile.txt"
all_text <- readLines(fname)
all_lines <- tibble(text = all_text)
tidiedtext <- all_lines %>%
  mutate(linenumber = row_number(),
         section_id = cumsum(str_detect(text, regex("^#", ignore_case = TRUE)))) %>%
  filter(!str_detect(text, regex("^#"))) %>%
  ungroup()

在tidiedtext 中为每行的相应节号添加一列。

是否可以在对mutate() 的调用中添加一行来添加这样的列?还是我应该使用另一种方法?

【问题讨论】:

    标签: r tidyverse tidytext


    【解决方案1】:

    我不希望你重写整个脚本,但我只是发现这个问题很有趣,并想添加一个基本的 R 暂定:

    parse_data <- function(file_name) {
      all_rows <- readLines(file_name)
      indices <- which(grepl('#', all_rows))
      splitter <- rep(indices, diff(c(indices, length(all_rows)+1)))
      lst <- split(all_rows, splitter)
      lst <- lapply(lst, function(x) {
        data.frame(section=x[1], value=x[-1], stringsAsFactors = F)
      })
      line_nums = seq_along(all_rows)[-indices]
      df <- do.call(rbind.data.frame, lst)
      cbind.data.frame(df, linenumber = line_nums)
    }
    

    使用名为 ipsum_data.txt 的文件进行测试:

    parse_data('ipsum_data.txt')
    

    产量:

     text                        section          linenumber
     Lorem ipsum dolor           # Section 1 Name 2         
     sit amet ... (et cetera)    # Section 1 Name 3         
     <multiple lines here again> # Section 2 Name 5   
    

    文件ipsum_data.txt 包含:

    # Section 1 Name
    Lorem ipsum dolor
    sit amet ... (et cetera)
    # Section 2 Name
    <multiple lines here again>
    

    我希望这证明有用。

    【讨论】:

    • 感谢您的回复。这很有帮助。重写脚本对我来说没什么大不了的,但我认为另一种解决方案在简洁性方面更符合我的要求。
    【解决方案2】:

    这是一种使用grepl 的方法,为简单起见,使用if_else 和tidyr::fill,但原来的方法没有任何问题;它与 tidytext book 中使用的非常相似。另请注意,添加行号后的过滤将使一些不存在。如果重要,请在filter 之后添加行号。

    library(tidyverse)
    
    text <- '# Section 1 Name
    Lorem ipsum dolor
    sit amet ... (et cetera)
    # Section 2 Name
    <multiple lines here again>'
    
    all_lines <- data_frame(text = read_lines(text))
    
    tidied <- all_lines %>% 
        mutate(line = row_number(),
               section = if_else(grepl('^#', text), text, NA_character_)) %>% 
      fill(section) %>% 
      filter(!grepl('^#', text))
    
    tidied
    #> # A tibble: 3 × 3
    #>                          text  line          section
    #>                         <chr> <int>            <chr>
    #> 1           Lorem ipsum dolor     2 # Section 1 Name
    #> 2    sit amet ... (et cetera)     3 # Section 1 Name
    #> 3 <multiple lines here again>     5 # Section 2 Name
    

    或者,如果您只想格式化已有的数字,只需将 section_name = paste('Category', section_id, 'Name') 添加到您的 mutate 调用中。

    【讨论】:

    • 谢谢!这几乎就是我想要的。
    猜你喜欢
    • 1970-01-01
    • 2015-02-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多