【问题标题】:Ordering bars in geom_segment based gantt chart using ggplot with duplicate y-factors使用具有重复 y 因子的 ggplot 在基于 geom_segment 的甘特图中排序条形图
【发布时间】:2015-11-26 15:36:01
【问题描述】:

问题

我有一个数据集se.df(问题底部的数据),我通过使用ggplot 和facet_grid 将其可视化为分解甘特图。但是,y 标签没有按照我对aes 指定的顺序进行排序

library(ggplot2)
base <- ggplot(
  se.df,
  aes(
    x = Start.Date, reorder(Action,Start.Date), color = Comms.Type
  ))
base + geom_segment(aes(
  xend = End.Date,ystart = Action, yend = Action
), size = 5) + 
facet_grid(Source ~ .,scale = "free_y",space = "free_y", drop = TRUE)

在这张详细图片中,您可以看到以下条形:

  1. Start.Date 订单中未显示
  2. Action 未订购。澄清一下,条形图应按 Start.Date 排序,然后按字母顺序按Action

如何根据Start.Date 和Action 对每个因子内的条进行排序?

更新

@heathobrien 提供了一个解决方案,解决了我通过Start.Date 订购酒吧的问题,而不是由重复因素引起的问题 - 这是我的实际数据所具有的。

Action 中有两个 "Inform colleges" 实例,这导致来自@heathobrien 的以下代码中的顺序错误,在图像中以红色虚线椭圆突出显示:

se.df <-se.df[order(se.df$Start.Date,se.df$Action),]
se.df$Action <- factor(se.df$Action, levels=unique(se.df$Action))
ggplot(se.df, aes(x = Start.Date, color = Comms.Type)) +
  geom_segment(aes(xend = End.Date, y = Action, yend = Action), size = 5) +
  facet_grid(Source ~ .,scale = "free_y",space = "free_y", drop = TRUE)

如何将此 data.frame 提供给 ggplot 以使每个 facet_grid 内的顺序保持一致?

更多细节

关于制作甘特图和排序因素有很多问题,我根据其他人的回答做出了一些决定:

  1. geom_segment

许多提问者都使用过geom_linerange,但无法使用coord_flip with non-cartesian coordinate systems。解决方案很复杂,我已经使用 geom_segment 缓解了这些问题。

  1. reorder 内aes

几乎规范的条形排序问题uses reorder。但是,这不适用于我的数据,即使使用 transform 而不是直接将 order 指定给 aes。我很乐意找到任何可行的解决方案。

数据

se.df <- structure(list(Source = structure(c(2L, 2L, 2L, 2L, 2L, 2L, 2L, 
2L, 2L, 2L, 2L, 2L, 2L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 
1L, 1L, 3L, 3L, 3L, 3L, 3L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 
2L, 2L, 2L), .Label = c("a", "b", "c"), class = c("ordered", 
"factor")), Action = structure(c(21L, 30L, 19L, 27L, 16L, 17L, 
18L, 13L, 12L, 3L, 1L, 8L, 4L, 21L, 20L, 27L, 15L, 17L, 18L, 
14L, 26L, 2L, 8L, 5L, 22L, 26L, 2L, 8L, 5L, 22L, 22L, 11L, 7L, 
24L, 29L, 6L, 23L, 25L, 25L, 10L, 28L, 9L), .Label = c("Add OA \"Act on Acceptance\" to websites", 
"Add RDM liaison presece to divisional and departmental websites", 
"All-staff message from VC and/or Pro-VC (Research)", "Arrange OA Briefing for every department", 
"Arrange RDM Briefing for every department", "Brief Communication Officers Network", 
"Brief Conference of Colleges", "Brief divisional board/commitees", 
"Brief Faculty IT Officers", "Brief Research Committee", "Brief Senior Tutors", 
"Brief/mobilise internal comms officers", "Brief/mobilise ORFN", 
"Brief/mobilise Subject Librarians", "Ceate template slides for colleagues to use in delivering RDM Briefings", 
"Create template slides for colleagues to use in delivering OA Briefings", 
"Create template text & icon for use on websites", "Draft material for use in staff induction", 
"Ensure webpages for ORA & Symplectic Elements  are updated & consistent", 
"Ensure webpages for ORA-Data are updated & consistent", "Finalise key messages and draft campaign text", 
"Inform colleges ", "Inform Heads of Departments and Research Directors", 
"Present at Departmental Administrator's Meeting", "Present at HAF meeting", 
"Present at UAS Conference", "Produce hard copy materials to promote message ", 
"Update Divisional Board", "Update Library Committee (CLIPS)", 
"Update OAO website content for HEFCE/REF"), class = "factor"), 
    Start.Date = structure(c(1435705200, 1435705200, 1438383600, 
    1441062000, 1441062000, 1441062000, 1441062000, 1444518000, 
    1444518000, 1425168000, 1420070400, 1444518000, 1444518000, 
    1441062000, 1441062000, 1441062000, 1441062000, 1441062000, 
    1441062000, 1438383600, 1441062000, 1420070400, 1444518000, 
    1444518000, 1443654000, 1441062000, 1420070400, 1444518000, 
    1444518000, 1443654000, 1441062000, 1444518000, 1449273600, 
    1444518000, 1444518000, 1445036400, 1441062000, 1443740400, 
    1443740400, 1443740400, 1447459200, 1443740400), class = c("POSIXct", 
    "POSIXt"), tzone = ""), End.Date = structure(c(1440975600, 
    1440975600, 1443567600, 1443567600, 1443567600, 1443567600, 
    1443567600, 1449273600, 1449273600, 1430348400, 1446249600, 
    1449273600, 1449273600, 1446249600, 1446249600, 1443567600, 
    1443567600, 1443567600, 1443567600, 1443567600, 1443567600, 
    1443567600, 1449014400, 1449014400, 1451520000, 1443567600, 
    1443567600, 1449014400, 1449014400, 1451520000, 1443567600, 
    1449014400, 1449619200, 1449014400, 1449014400, 1446249600, 
    1446249600, 1449792000, 1449792000, 1449792000, 1447804800, 
    1449792000), class = c("POSIXct", "POSIXt"), tzone = ""), 
    Comms.Type = structure(c(3L, 7L, 7L, 6L, 5L, 7L, 8L, 4L, 
    4L, 2L, 7L, 1L, 1L, 3L, 7L, 6L, 5L, 7L, 8L, 4L, 5L, 7L, 1L, 
    1L, 5L, 5L, 7L, 1L, 1L, 5L, 1L, 1L, 1L, 5L, 3L, 1L, 1L, 5L, 
    5L, 1L, 1L, 1L), .Label = c("Briefing", "Email", "Mixed Media", 
    "Mobilisation", "Presentations", "Printed Materials", "Website", 
    "Workshop"), class = "factor")), .Names = c("Source", "Action", 
"Start.Date", "End.Date", "Comms.Type"), row.names = c(NA, -42L
), class = c("tbl_df", "tbl", "data.frame"))

【问题讨论】:

  • 我认为,如果您在绘图之前按 start.date 对数据框进行排序,那么您将得到解决
  • 澄清一下:您想按字母顺序或按 Comms.Type 排序 Action?
  • @Felix 按字母顺序是目标,很抱歉省略了这一点。 heathobrien 我会在回到我的机器时检查一下
  • @heathobrien 以下se.df &lt;-se.df[order(se.df$Start.Date,se.df$Action),] 无法解决我的图像中显示的错误排序。通过此重新排序从aes 中删除reorder(Action,Start.Date 行确实会按字母顺序重新排序Action,但不会在ggplot 输出中按Start.Date 排序。

标签: r ggplot2 dataframe visualization


【解决方案1】:

一旦你按照你想要的顺序对数据框进行了排序,你应该能够将它用作你的因子的级别:

se.df <-se.df[order(se.df$Start.Date,se.df$Action),]
se.df$Action <- factor(se.df$Action, levels=unique(se.df$Action))
ggplot(se.df, aes(x = Start.Date, color = Comms.Type)) +
  geom_segment(aes(xend = End.Date, y = Action, yend = Action), size = 5) +
  facet_grid(Source ~ .,scale = "free_y",space = "free_y", drop = TRUE)

【讨论】:

  • 感谢您的回答,很遗憾Source“b”中仍然存在这两个问题。 1)“确保网页”项目显示在“完成关键消息”栏之前,尽管开始日期较晚,2)许多活动于 2015 年 9 月 1 日开始,但奇怪的是“通知学院”栏出现在“制作硬拷贝”,尽管它之前是按字母顺序排列的。对不起,我的数据有这么长的整体,让人难以谈论:(
  • ... 在我的机器上评估更新的答案之前,我写道,我现在有了我需要的订单 - 谢谢!要不要我再给你上传一张图片,省去你的麻烦?
  • 这些问题是因为您在其中多次使用相同的操作名称但开始时间不同。我不确定解决方案是什么(除了重命名它们),因为你不能有两个具有相同名称的因子水平
  • 当然。这将是最有帮助的
  • 感谢您的回答 - 在我可以保证唯一条目的情况下它会很有用。 @C8H10N4O2 答案能够克服重复问题,并迫使我学习这种 %>% 的精神错乱。
【解决方案2】:

我认为这就是 OP 正在寻找的:

我必须创建一个合成的taskID 来传递订单(通过增加 Start.Date,通过 Action 增加 alpha)。顺便说一句,如果你想按Action的字母顺序排列,你需要改变因子的顺序或者转换成char。

# first let's order the DF the way we want it to appear 
#    (higher taskID's first)

# dplyr-free version
se.df$Action <- as.character(se.df$Action)  
se.df <- se.df[order(se.df$Start.Date, se.df$Action), ]
se.df$taskID <- as.factor(nrow(se.df):1)


library(ggplot2)
ggplot(se.df, aes(x = Start.Date, y=taskID, color = Comms.Type)) +
  scale_y_discrete(breaks=se.df$taskID, labels = se.df$Action) + 
  geom_segment(aes(xend = End.Date, y = taskID, yend = taskID), size = 5) +
  facet_grid(Source ~ .,scale = "free_y",space = "free_y", drop = TRUE)

【讨论】:

  • 太好了!你处理了有问题的“通知大学”副本,谢谢。我有很多东西要离开去学习(%&gt;% syntax and dplyr` 包对我来说是新的),但这是我要做的事情:) 我不明白为什么我的问题被否决了。我觉得我的例子是最小的和可重复的,我整理的资源足以让其他遇到类似问题的人。无论如何,感谢您的宝贵时间。
  • @MartinJohnHadley 仅将其更改为基本 R。请注意,dplyr::arrange 使用 C++ 进行字母排序,因此使用order,“通知学院”现在出现在“通知负责人...”之前,Dplyr 在这种情况下并没有很大的帮助,但在很多情况下都救了我。
【解决方案3】:

这个问题是在 2015 年提出的,在 tidyverse 和优秀的 forcats 库出现之前。

这里有一个tidyverse 解决问题的方法:

使用arrange将数据按Start.Date和Action排序,然后使用row_number()创建task_id

library("tidyverse")
se.df <- se.df %>%
  arrange(desc(Start.Date), Action) %>%
  mutate(task_id = row_number())

使用fct_reorder 将Action 转换为由task_id 排序的因子。

se.df <- se.df %>%
  mutate(
    Action = fct_reorder(Action, task_id),
    Action = fct_rev(Action)
  )

现在我们可以绘制这些数据而无需替换坐标轴标签:

se.df %>%
  ggplot(aes(x = Start.Date, y = Action, color = Comms.Type)) +
  geom_segment(aes(xend = End.Date, y = Action, yend = Action), size = 5) +
  facet_grid(Source ~ ., scale = "free_y", space = "free_y", drop = TRUE)

【讨论】:

    猜你喜欢
    • 2021-03-04
    • 1970-01-01
    • 2019-03-10
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2023-01-19
    • 1970-01-01
    相关资源
    最近更新 更多