【发布时间】:2014-03-20 20:39:57
【问题描述】:
我是一名社会科学研究人员,致力于想象人们如何随着时间的推移在社区中扮演不同的角色。
我已将人们每月的行为归类为角色类别,现在我想可视化每个(相对)时间段内每个角色的人数和比例。
现在,数据保存在 CSV 格式中,如下所示:
ID T1 T2 T3 ...
1 2 2 3
2 1 0 2
3 1 2 1
...
其中 X(ij) 是我在第 j 个月所处的集群 ID。
我想要的是这样的(我在 LibreOffice 中创建的)。
我相信我需要使用 ggplot2,但我一直在努力弄清楚如何以 ggplot 喜欢的格式获取数据。
我想我的第一个任务是总结每个时间段的每个集群?有没有简单的方法可以做到这一点?
我可以用下面的代码做到这一点,但是它很糟糕而且很乱,一定有更好的方法吗?
clus1 <- apply(clusters, 2, function(x) {sum(x=='1', na.rm=TRUE)})
clus2 <- apply(clusters, 2, function(x) {sum(x=='2', na.rm=TRUE)})
clus3 <- apply(clusters, 2, function(x) {sum(x=='3', na.rm=TRUE)})
clus0 <- apply(clusters, 2, function(x) {sum(x=='0', na.rm=TRUE)})
clusters2 <- data.frame(clus0, clus1, clus2, clus3)
c2 <- t(clusters2)
c3 <- as.data.frame(c2)
c3$id = c('Low Activity Cluster', 'Cluster 1', 'Cluster 2', 'Cluster 3')
c3 <- c3[order(c3$'id'),]
print(ggplot(melt(c3, id.vars="id")) +
geom_area(aes(x=variable, y=value, fill=id, group=id), position="fill"))
这会导致示例数据如下所示:
id T1 T2 T3
Low Activity Cluster 0 1 0
Cluster 1 2 0 1
Cluster 2 1 2 1
Cluster 3 0 0 1
这是正确的策略吗?
【问题讨论】: