【问题标题】:Produce Bar Chart of Filtered Columns with ggplot2使用 ggplot2 生成过滤列的条形图
【发布时间】:2020-08-02 15:39:00
【问题描述】:

您能告诉我如何制作如图所示的图表吗? 我只想选择每个城市的前 2 个社区(基于房价中位数的前 2 个社区)并显示它们的中位数价格。当然,如果酒吧是不同的颜色更好.. (请注意,我是手动生成中间价并在 Excel 中绘制的,因此它们不代表真实值)

    glimpse(CityNeighbourhoodPrice)
Observations: 37,245
Variables: 3
$ City          <fct> Amsterdam, Amsterdam, Amsterdam...
$ Neighbourhood <fct> A,B,C,D,E,F,G,H,I,J,K...
$ Price         <int> 970, 1320, 2060, 2480, 1070, 12...

这是我目前的代码(不起作用):

CityNeighbourhoodPrice %>% 
  group_by(Neighbourhood) %>%
  count(n) %>%
  top_n(2, MedPrice) %>%
  summarise(MedPrice = median(Price, na.rm = TRUE)) %>%
  ggplot(aes(x = reorder(Neighbourhood,-MedPrice), y = MedPrice)) +
  geom_col(fill = "tomato3", width = 0.5)+
  labs(title="Ordered Bar Chart", 
       subtitle="Average Price by each Property Type", 
       caption="Image: 5") + 
  theme(axis.text.x = element_text(angle=65, vjust=0.6))

【问题讨论】:

  • 请在您的问题中包含数据集或数据集的子集。看起来您只需要来自 City、Neighborhood 和 Price 的数据。或许可以解释一下字母 A、B、C... 的含义以及它们的来源。
  • 字母 A、B、C... 是什么意思,看起来好像它们反映了某种“顶级”的衡量标准——这是如何定义的或者是什么变量捕获了这个属性? (抱歉只允许 5 分钟编辑评论!)
  • 嗨,对不起。 A,B,C 指的是邻里!我相信问题中提供了数据集?我只想选前 2 个街区。前 2 个街区是根据 MEDIAN PRICE 确定的。我会澄清这个问题@Peter
  • 谢谢。您提供了数据集的视图,但复制和粘贴并不容易,因此可用于回答问题。您能否提供一个摘录,也许使用 dput() 以数据框的形式,仅包括适用于您的问题的变量。如果您不确定,请参阅 minimal reproducible example 了解如何执行此操作的示例。这会让其他人更容易帮助你。
  • 抱歉,我不确定如何使用 dput(),我已经编辑了我的问题以仅包含所需的变量..

标签: r ggplot2 dplyr data-visualization summarize


【解决方案1】:

使用一些随机示例数据,试试这个:

# Example data
set.seed(42)

CityNeighbourhoodPrice <- data.frame(
  City = rep(c("Amsterdam", "Berlin", "Edinburgh"), each = 30),
  Neighbourhood = rep(sample(LETTERS[1:5], 30, replace = TRUE), 3),
  Price = 3000 * runif(3 * 30)
)

library(ggplot2)
library(dplyr)
library(forcats)

# Plot
CityNeighbourhoodPrice %>% 
  group_by(City, Neighbourhood) %>%
  summarise(MedPrice = median(Price, na.rm = TRUE)) %>%
  top_n(2, MedPrice) %>%
  ungroup() %>% 
  arrange(City, MedPrice) %>% 
  mutate(City_Neighbourhood = paste0(Neighbourhood, "\n", City),
         City_Neighbourhood = forcats::fct_inorder(City_Neighbourhood)) %>% 
  ggplot(aes(x = City_Neighbourhood, y = MedPrice)) +
  geom_col(fill = "tomato3", width = 0.5)+
  labs(title="Ordered Bar Chart", 
       subtitle="Average Price by each Property Type", 
       caption="Image: 5") + 
  theme(axis.text.x = element_text(angle=65, vjust=0.6))

由reprex package (v0.3.0) 于 2020 年 4 月 20 日创建

【讨论】:

    【解决方案2】:

    另一种解决方案可能是:

    假设您的数据如下所示:

    library(dplyr)
    library(ggplot)
    
    data <- data.frame(Price=c(970, 245, 564, 895, 431, 100), City=c("Amsterdam", "Athens", "Amsterdam", "London", "Berlin", "Netherlands"), Neighborhood=c("A", "B", "D", "C", "E", "F"))
    

    然后你做:

    example_plot <- data %>%
      select(Price, City, Neighborhood) %>%
      group_by(City) %>%
      top_n(., 2, wt=Price) %>%
      spread(Neighborhood, Price) %>%
      data.frame %>%
      mutate(., Average=rowMeans(.[,-1], na.rm = TRUE)) %>%
      ggplot(., aes(City, Average, fill=City)) +
      ggtitle(str_wrap(c("Median Price for the Top-2 Neighborhoods in Different Cities:"), 20)) +
      theme_fivethirtyeight() +
      theme(legend.position = "none", plot.title = element_text(size= 22), axis.text = element_text(size=14))+
      geom_bar(stat = "identity") +
      geom_text(aes(x = City, y = Average, label = Average ), colour = "white", size = 11, vjust=1.2)
    

    它给你:

    【讨论】:

    • 对不起,Stefan 提供的输出与我的预期最相似,因为我希望每个城市被分成两个社区。非常感谢您的努力!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-09-27
    • 2021-04-24
    • 2016-07-15
    • 1970-01-01
    相关资源
    最近更新 更多