【发布时间】:2018-04-21 02:05:47
【问题描述】:
如何使用 tm 包在 R 中打印一个小样本或第一行语料库?我有一个非常大的语料库(> 1 GB)并且正在做一些文本清理。我想在应用清洁程序时进行测试。只打印语料库的第一行或前几行是理想的。
# Load Libraries
library(tm)
# Read in Corpus
corp <- SimpleCorpus( DirSource(
"C:/TextDocument"))
# Remove puncuation
corp <- removePunctuation(corp,
preserve_intra_word_contractions = TRUE,
preserve_intra_word_dashes = TRUE)
我尝试了几种访问语料库的方法:
# Print first line of first element of corpus
corp[[1]][[1]]
# Print first line using 'content' element of corpus
corp[[1]]$content[[1]]
这两种情况都会导致运行时间过长而没有所需的输出。
tm包中的原始语料可以用于示例目的。
data("crude")
【问题讨论】:
-
你为什么不先获取你的语料库的一个子集,对它进行所有的文本清理测试,然后在完整的语料库上做呢?或切换到量子。并行工作。从语料库中获取信息的最快方法也是 corp[[1]]$content[[1]]。你可以用微基准做一些测试来检查。
标签: r text-mining tm corpus