【问题标题】:How to transform a Document Term Matrix in R?如何在 R 中转换文档术语矩阵?
【发布时间】:2018-11-26 14:25:06
【问题描述】:

您好,我有一个文档术语矩阵,我使用tidy() 函数对其进行了转换,它运行良好。我想根据单词的频率绘制一个词云。所以我的转换表如下所示:

> head(Wcloud.Data)
# A tibble: 6 x 3
  document term       count
  <chr>    <chr>      <dbl>
1 1        accept         1
2 1        access         1
3 1        accomplish     1
4 1        account        4
5 1        accur          2
6 1        achiev         1

我有 33,647,383 个观察值,因此它是一个非常大的数据框。如果我使用max() 函数,我会得到一个非常高的数字(64116),但我的数据框中没有一个词的频率为 64116。此外,如果我用 wordcloud() 绘制闪亮的数据框,它会多次绘制相同的词。此外,如果我想对我的列进行排序 count 它不起作用 - sort(Wcloud.Data$count,decreasing = TRUE)。所以有些事情是不正确的,但我不知道,什么以及如何解决它。有人知道吗?

这是我的文档术语矩阵的摘要,在将其转换为数据框之前:

> observations.tf
<<DocumentTermMatrix (documents: 76717, terms: 4234)>>
Non-/sparse entries: 33647383/291172395
Sparsity           : 90%
Maximal term length: 15
Weighting          : term frequency (tf)

更新:我添加了我的数据框的图片

【问题讨论】:

  • 您能否向我们提供数据子集Wcloud.Data(可能使用dput),以便我们可以在您的数据集上重现问题?我想我有一个解决方案给你,但需要在当地确认。谢谢:)
  • 同一个词出现是正常的,因为您有多个文档 (76717),如果一个词出现在多个文档中的频率很高,它会被打印多次。如果您想要仅包含单词的 wordcloud,请删除文档并汇总每个单词的数字。
  • @phiver 感谢您的回答。我怎样才能自动解决这个问题?我不希望它有多个。
  • @mysteRious 我不知道为什么,但输出 dput 有问题。或者 R 正在计算,它需要一些时间。你的想法是什么?
  • 任何可以使用 100-1000 行 Wcloud.Data 的东西都会有所帮助。

标签: r dataframe dataset transformation word-cloud


【解决方案1】:

使用dplyr您可以如下:

library("tm")
library("SnowballC")
library("wordcloud")
library("RColorBrewer")

Wcloud.Data<- data.frame(Document= c(rep(1,6)), 
                         term = c("accept", "access","accomplish", "account", "accur", "achiev"),
                         count = c(1,1,1,4,2,1))

Data<-Wcloud.Data %>% 
  group_by(term) %>% 
  summarise(Frequency = sum(count))
set.seed(1234)
wordcloud(words = Data$term, freq = Data$Frequency, min.freq = 1,
          max.words=200, random.order=FALSE, rot.per=0.35, 
          colors=brewer.pal(8, "Dark2"))

在另一边,图书馆quanteda和tibble可以帮助您重写术语频率矩阵。我会把你的一个例子与它合作:

library(tibble)
library(quanteda)
Data <- data_frame(text = c("Chinese Beijing Chinese",
                              "Chinese Chinese Shanghai",
                              "this is china",
                              "china is here",
                              'hello china',
                              "Chinese Beijing Chinese",
                              "Chinese Chinese Shanghai",
                              "this is china",
                              "china is here",
                              'hello china',
                              "Kyoto Japan",
                              "Tokyo Japan Chinese",
                              "Kyoto Japan",
                              "Tokyo Japan Chinese",
                              "Kyoto Japan",
                              "Tokyo Japan Chinese",
                              "Kyoto Japan",
                              "Tokyo Japan Chinese",
                              'japan'))
DocTerm <- quanteda::dfm(Data$text)
DocTerm
# Document-feature matrix of: 19 documents, 11 features (78.5% sparse).
# 19 x 11 sparse Matrix of class "dfm"
# features
# docs     chinese beijing shanghai this is china here hello kyoto japan tokyo
# text1        2       1        0    0  0     0    0     0     0     0     0
# text2        2       0        1    0  0     0    0     0     0     0     0
# text3        0       0        0    1  1     1    0     0     0     0     0
# text4        0       0        0    0  1     1    1     0     0     0     0
# text5        0       0        0    0  0     1    0     1     0     0     0
# text6        2       1        0    0  0     0    0     0     0     0     0
# text7        2       0        1    0  0     0    0     0     0     0     0
# text8        0       0        0    1  1     1    0     0     0     0     0
# text9        0       0        0    0  1     1    1     0     0     0     0
# text10       0       0        0    0  0     1    0     1     0     0     0
# text11       0       0        0    0  0     0    0     0     1     1     0
# text12       1       0        0    0  0     0    0     0     0     1     1
# text13       0       0        0    0  0     0    0     0     1     1     0
# text14       1       0        0    0  0     0    0     0     0     1     1
# text15       0       0        0    0  0     0    0     0     1     1     0
# text16       1       0        0    0  0     0    0     0     0     1     1
# text17       0       0        0    0  0     0    0     0     1     1     0
# text18       1       0        0    0  0     0    0     0     0     1     1
# text19       0       0        0    0  0     0    0     0     0     1     0

Mat<-quanteda::convert(DocTerm,"data.frame")[,2:ncol(DocTerm)] # Converting to a Dataframe without taking into account the text variable
Result<- colSums(Mat) # This is what you are interested in
names(Result)<-colnames(Mat)
# > Result
# chinese  beijing shanghai     this       is    china     here    hello    kyoto    japan 
# 24        4        4        4        8       12        4        4        8       18 

【讨论】:

  • 谢谢卡尔斯,但我的问题是不同的。 Phiver发现我的问题到了。 span>
  • 好的,现在我编辑了我的问题来获得答案。此外,如果您只需尝试通过“术语”组成的总和(count)组,它应该解决。 span>
  • 那是我的问题:出现同一个词是正常的,因为您有多个文档 (76717),如果一个词以高频率出现在多个文档中,它将被打印多次。
  • 是,即术语频率矩阵是什么。每份文件的单词数。如果最典型的单词发生在多个文档中,则它看起来多次。 span>
  • 所以我能做什么? span>
猜你喜欢
  • 1970-01-01
  • 2021-07-27
  • 2018-05-04
  • 1970-01-01
  • 2015-05-19
  • 1970-01-01
  • 2018-06-24
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多