【问题标题】:R: Convert a "Term Document Matrix" to a "Corpus"R:将“术语文档矩阵”转换为“语料库”
【发布时间】:2021-07-27 09:20:02
【问题描述】:

我正在使用 R 编程语言。我正在尝试按照本教程的说明(https://cran.r-project.org/web/packages/tidytext/vignettes/tidying_casting.html)学习如何将“术语文档矩阵”转换为“语料库”。但是,我不清楚本教程中提供的解释,我不知道该怎么做。

使用公开的莎士比亚戏剧,我创建了术语文档矩阵,如下所示:

#load libraries
library(dplyr)
library(pdftools)
library(tidytext)
library(textrank)
library(tm)

#1st document
url <- "https://shakespeare.folger.edu/downloads/pdf/hamlet_PDF_FolgerShakespeare.pdf"

article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
  unnest_tokens(sentence, text, token = "sentences") %>%
  mutate(sentence_id = row_number()) %>%
  select(sentence_id, sentence)


article_words <- article_sentences %>%
  unnest_tokens(word, sentence)


article_words_1 <- article_words %>%
  anti_join(stop_words, by = "word")

#2nd document
url <- "https://shakespeare.folger.edu/downloads/pdf/macbeth_PDF_FolgerShakespeare.pdf"

article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
  unnest_tokens(sentence, text, token = "sentences") %>%
  mutate(sentence_id = row_number()) %>%
  select(sentence_id, sentence)


article_words <- article_sentences %>%
  unnest_tokens(word, sentence)


article_words_2<- article_words %>%
  anti_join(stop_words, by = "word")


#3rd document
url <- "https://shakespeare.folger.edu/downloads/pdf/othello_PDF_FolgerShakespeare.pdf"

article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
  unnest_tokens(sentence, text, token = "sentences") %>%
  mutate(sentence_id = row_number()) %>%
  select(sentence_id, sentence)


article_words <- article_sentences %>%
  unnest_tokens(word, sentence)


article_words_3 <- article_words %>%
  anti_join(stop_words, by = "word")

从这里,我创建了实际的“术语文档矩阵”:

library(tm)

#create term document matrix
tdm <- TermDocumentMatrix(Corpus(VectorSource(rbind(article_words_1, article_words_2, article_words_3))))

#inspect the "term document matrix" (I don't know why this is producing an error)
inspect(tdm)

现在,我不确定如何使用本教程 (https://cran.r-project.org/web/packages/tidytext/vignettes/tidying_casting.html) 中的说明并将“术语文档矩阵”转换为“语料库”。

这是应该的吗?

library(quanteda)
d <- quanteda::dfm(tdm, verbose = FALSE)

谁能告诉我如何解决这个问题?

谢谢

【问题讨论】:

    标签: r text nlp text-mining


    【解决方案1】:

    您可以尝试quanteda::corpus 将数据框转为语料库。

    data <- rbind(article_words_1, article_words_2, article_words_3)
    corp <- quanteda::corpus(data, text_field = 'word')
    

    【讨论】:

    • 谢谢!但是有没有办法直接将“tdm”转换成“语料库”呢?
    • 很遗憾,我找不到这样做的方法。
    【解决方案2】:

    回答您在 Ronak 答案中的问题。

    您无法将 tdm 转换为语料库,因为 tdm 已经在文档中汇总了字数,并且您丢失了句子的顺序。使用 quanteda,您可以在 dfm 上执行多项操作,例如替换单词或删除停用词。请参阅基于莎士比亚第一个文本的示例以显示该过程。

    library(dplyr)
    library(tidytext)
    library(pdftools)
    
    library(quanteda)
    
    url <- "https://shakespeare.folger.edu/downloads/pdf/hamlet_PDF_FolgerShakespeare.pdf"
    
    article <- pdf_text(url)
    article_sentences <- tibble(text = article) %>%
      unnest_tokens(sentence, text, token = "sentences") %>%
      mutate(sentence_id = row_number()) %>%
      select(sentence_id, sentence)
    
    # added count to simulate a dtm
    article_words <- article_sentences %>%
      unnest_tokens(word, sentence) %>% 
      group_by(sentence_id, word) %>% 
      summarise(count = n())
    
    # cast into a quanteda dfm
    my_dfm <- cast_dfm(article_words, sentence_id, word, count)
    
    # using quanteda's dfm_remove to remove stopwords from a dfm.
    my_dfm <- dfm_remove(my_dfm, stopwords())
    

    使用dfm_replace,您可以替换某些可能有错误和/或您不想使用的标点符号的单词。之后,您可以使用 dfm_compress 将具有相同名称的特征组合成 1 个特征。

    但您最好尝试获取原始数据,而不是从 tdm 开始。

    【讨论】:

    • 感谢您的回复 - 问题是,对于我的真实数据“:我只有“tdm”文件,我不再有“article_1”、“article_2”和“article_3”的等价物“。我正在尝试将“tdm”文件转换为“语料库”,以便我可以使用 quanteda 库中的某些函数。这仍然可能吗?
    • stackoverflow.com/questions/67376045/… :在这里,我正在尝试删除“停用词”和“标记化”(都使用“quanteda”)“tdm”文件,但我不能这样做,因为 quanteda不接受“tdm”文件。假设您已经有了“tdm”文件——有没有办法“标记化”和“删除停用词”?谢谢
    • @Noob,如果你已经有一个 tdm,你可以在 tdm 对象上使用as.dfm。这会将 tdm 转换为可以与 quanteda 一起使用的 dfm。
    • 哇,谢谢你的建议!我一定会尝试的!
    • @Noob,您无法将 dfm 转换为令牌对象。你需要使用b %&gt;% dfm_remove(stopwords("en"))
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2018-05-04
    • 2022-01-21
    • 1970-01-01
    • 2018-11-26
    • 2017-10-16
    • 1970-01-01
    • 2018-06-24
    相关资源
    最近更新 更多