【问题标题】:R Help! Calculate the proportion per subgroup救命!计算每个子组的比例
【发布时间】:2021-02-15 02:04:18
【问题描述】:

我有以下 数据集,称为 GrossExp3,涵盖 15 个报告国从(1998 年至 2018 年)到所有年份的双边出口(以 1000 美元为单位)可用的伙伴国家 它涵盖以下四个变量: Year, ReporterName (= exporter) , PartnerName (= export destination), 'TradeValue in 1000 USD' (= 出口到目的地的价值) PartnerName 列还包括一个名为“All”的条目,它是记者每年所有出口的总和

这是我的数据头

> head(GrossExp3, n = 20)
    Year ReporterName          PartnerName TradeValue in 1000 USD
 1: 2018       Angola          Afghanistan                 19.353
 2: 2018       Angola              Albania                  2.380
 3: 2018       Angola              Andorra                  0.326
 4: 2018       Angola United Arab Emirates             884725.078
 5: 2018       Angola            Argentina                 61.362
 6: 2018       Angola              Armenia                 60.105
 7: 2018       Angola       American Samoa                 12.007
 8: 2018       Angola  Antigua and Barbuda                422.006
 9: 2018       Angola            Australia              40220.092
10: 2018       Angola              Austria                433.699

这是我的数据摘要

> summary(GrossExp3)
      Year      ReporterName       PartnerName        TradeValue in 1000 USD
 Min.   :1998   Length:37398       Length:37398       Min.   :       0      
 1st Qu.:2004   Class :character   Class :character   1st Qu.:      39      
 Median :2009   Mode  :character   Mode  :character   Median :     596      
 Mean   :2009                                         Mean   :  135605      
 3rd Qu.:2014                                         3rd Qu.:   10209      
 Max.   :2019                                         Max.   :47471515 

我的目标是按年按总出口百分比(所有百分比得分均超过 1%)过滤每个国家最重要的出口目的地,并跟踪其随时间的变化情况。 我特别想

  1. 添加一个名为“百分比”的附加列,其中包含按年份划分的总出口百分比(按年份划分的“TradeValue in 1000 USD”中所有条目的总和)
  2. 删除所有百分比
  3. 按年份汇总每个 ReporterName 的数据
  4. 按百分比值降序排列

我尝试了什么 到目前为止,我通过首先过滤一个 ReporterName 和一年来尝试它

ONE_country <- GrossExp3 %>%
  group_by(Year, ReporterName) %>%
  filter(ReporterName == "Botswana", PartnerName != "All", Year == 2018) %>%
  arrange(desc(`TradeValue in 1000 USD`)) %>%
  summarize(Year, ReporterName, PartnerName, Percent = `TradeValue in 1000 USD`/sum(`TradeValue in 1000 USD`)*100)
head(ONE_country, n = 10)

我不确定我在这里得到的结果是否正确。 此外,我希望所有国家和年份的信息都保留在同一数据集中。 此外,我没有设法降低所有百分比 > 1,并且希望百分比在逗号后没有条目

另一个问题是,如果我不在函数中写,为什么 summarize 函数 不会返回所有列?

由于我在周末遇到了这些问题,我将非常感谢有关如何解决问题的任何建议! 一切顺利, 我喜欢

【问题讨论】:

    标签: r dplyr percentage conditional-formatting summarize


    【解决方案1】:

    如果没有可重复的数据,这很难回答,但这可能会有所帮助:

    GrossExp3 %>%
      group_by(Year, ReporterName) %>%
      add_tally(wt = `TradeValue in 1000 USD`, name = "TotalValue") %>%
      mutate(Percentage = 100 * (`TradeValue in 1000 USD` / TotalValue)) %>%
      filter(Percentage >= 1) %>%
      arrange(ReporterName, Year, desc(Percentage))
    

    我们使用add_tally() 按国家/地区按年份计算贸易总值,然后使用mutate() 计算每一行占该总值的百分比。然后我们可以排除百分比

    根据您上面提供的非常有限的 sn-p,这是返回的内容:

    # A tibble: 2 x 6
    # Groups:   Year, ReporterName [1]
       Year ReporterName PartnerName          `TradeValue in 1000 USD` TotalValue Percentage
      <dbl> <chr>        <chr>                                   <dbl>      <dbl>      <dbl>
    1  2018 Angola       United Arab Emirates                  884725.    925956.      95.5 
    2  2018 Angola       Australia                              40220.    925956.       4.34
    

    【讨论】:

    • 亲爱的 Semaphorism, 非常感谢您的快速回答和详细的解决方案!该代码运行良好!非常感谢!
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-10-29
    • 2021-01-05
    • 2021-06-20
    • 1970-01-01
    相关资源
    最近更新 更多