【发布时间】:2021-07-27 09:20:02
【问题描述】:
我正在使用 R 编程语言。我正在尝试按照本教程的说明(https://cran.r-project.org/web/packages/tidytext/vignettes/tidying_casting.html)学习如何将“术语文档矩阵”转换为“语料库”。但是,我不清楚本教程中提供的解释,我不知道该怎么做。
使用公开的莎士比亚戏剧,我创建了术语文档矩阵,如下所示:
#load libraries
library(dplyr)
library(pdftools)
library(tidytext)
library(textrank)
library(tm)
#1st document
url <- "https://shakespeare.folger.edu/downloads/pdf/hamlet_PDF_FolgerShakespeare.pdf"
article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
unnest_tokens(sentence, text, token = "sentences") %>%
mutate(sentence_id = row_number()) %>%
select(sentence_id, sentence)
article_words <- article_sentences %>%
unnest_tokens(word, sentence)
article_words_1 <- article_words %>%
anti_join(stop_words, by = "word")
#2nd document
url <- "https://shakespeare.folger.edu/downloads/pdf/macbeth_PDF_FolgerShakespeare.pdf"
article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
unnest_tokens(sentence, text, token = "sentences") %>%
mutate(sentence_id = row_number()) %>%
select(sentence_id, sentence)
article_words <- article_sentences %>%
unnest_tokens(word, sentence)
article_words_2<- article_words %>%
anti_join(stop_words, by = "word")
#3rd document
url <- "https://shakespeare.folger.edu/downloads/pdf/othello_PDF_FolgerShakespeare.pdf"
article <- pdf_text(url)
article_sentences <- tibble(text = article) %>%
unnest_tokens(sentence, text, token = "sentences") %>%
mutate(sentence_id = row_number()) %>%
select(sentence_id, sentence)
article_words <- article_sentences %>%
unnest_tokens(word, sentence)
article_words_3 <- article_words %>%
anti_join(stop_words, by = "word")
从这里,我创建了实际的“术语文档矩阵”:
library(tm)
#create term document matrix
tdm <- TermDocumentMatrix(Corpus(VectorSource(rbind(article_words_1, article_words_2, article_words_3))))
#inspect the "term document matrix" (I don't know why this is producing an error)
inspect(tdm)
现在,我不确定如何使用本教程 (https://cran.r-project.org/web/packages/tidytext/vignettes/tidying_casting.html) 中的说明并将“术语文档矩阵”转换为“语料库”。
这是应该的吗?
library(quanteda)
d <- quanteda::dfm(tdm, verbose = FALSE)
谁能告诉我如何解决这个问题?
谢谢
【问题讨论】:
标签: r text nlp text-mining