【问题标题】:Dataframe Groupby to combine multiple rows, summing float-type columnsDataframe Groupby 组合多行,对浮点型列求和
【发布时间】:2020-02-16 18:33:06
【问题描述】:

糟糕的标题,但在这里。我有一个 13,000 x 91 的数据框。 26 列是数字。这些行是单个项目,项目绩效按年份划分。像这样:

| Year | Control | Description   | USD_Cost | USD_Profit |
|------|---------|---------------|----------|------------|
| 1991 | A1      | A description | 1        | 2          |
| 1992 | A1      | A Description | 100      | 300        |
| 1991 | B1      | B Description | 3        | 50         |
| 1995 | C1      | C Description | 5        | 10         |
| 1990 | D1      | D Description | 2        | 1          |
| 1996 | D1      | D Description | 1        | 1          |

我不想记录每个项目在每个特定年份的表现,我只想记录每个项目持续了多长时间,以及项目的整体表现:

| Years | Control | Description   | USD_Cost | USDProfit |
|-------|---------|---------------|----------|-----------|
| 2     | A1      | A description | 101      | 302       |
| 1     | B1      | B Description | 3        | 50        |
| 1     | C1      | C Description | 5        | 10        |
| 2     | D1      | D Description | 3        | 2         |

Control 和 Description 不会更改,但以 USD 开头的数字列会跨行求和。并且有 26 美元列用于不同的性能方面。大约有 8000 个唯一的 Control Number,但 Year-ControlNumber 组合共有 13000 个。

我知道如何为一个元素分组(例如print(dft.groupby(['Control'])['USD_Cost', 'USD_Profit'].sum() ),但是当我这样做时,我想我会丢失所有非数字列。另外,我想避免输入所有 26 美元列的名称.

这可以用 groupby 完成吗?

【问题讨论】:

    标签: python pandas dataframe pandas-groupby


    【解决方案1】:

    所以我的解决方案是按“Control”分组,然后对每个组应用一个函数,该函数从第一行获取所有非数字数据(我假设非数字数据的所有行都相同) ,但取所有数值数据的总和。由于您不想将年数相加,因此将单独处理年数。

    我的代码:

    import pandas as pd
    import numpy as np
    
    
    def sum_project(project):
        # Since only numeric data and years are different,
        # we just take the first row
        project_summed = project.iloc[0, :]
    
        # sum all numeric data but exclude "Year"
        cols_numeric = project.select_dtypes([np.number]).columns
        cols_numeric = cols_numeric.drop(["Year"])
        project_summed[cols_numeric] = project[cols_numeric].sum()
    
        # Get year number
        project_summed["Year"] = len(project)
    
        return project_summed
    
    
    df = pd.DataFrame({
        "Year": [1991, 1992, 1991, 1995, 1990, 1996],
        "Control": ["A1", "A1", "B1", "C1", "D1", "D1"],
        "Description": [
            "A description",
            "A description",
            "B description",
            "C description",
            "D description",
            "D description"
        ],
        "USD_Cost": [1, 100, 3, 5, 2, 1],
        "USD_Profit": [2, 300, 50, 10, 1, 1],
    })
    
    findal_df = df.groupby(["Control"]).apply(sum_project)
    

    这给出了 final_df:

             Year Control    Description  USD_Cost  USD_Profit
    Control                                                   
    A1          2      A1  A description       101         302
    B1          1      B1  B description         3          50
    C1          1      C1  C description         5          10
    D1          2      D1  D description         3           2
    

    【讨论】:

    • 这很有道理 - 我今晚试试这个并报告。感谢您的指导!
    • 这很好用!谢谢你。我所做的一项调整是 cols_numeric = df.select_dtypes([np.number]).columns 可以概括为 cols_numeric = project.select_dtypes([np.number]).columns
    【解决方案2】:

    我认为这应该适合你

    columns = list(filter(lambda x: 'USD' in x, df.columns))
    df.groupby(['Control', 'Description'])[columns].sum()

    这将为您带来按控制、描述分组的所有列。这对你的工作来说不是问题,所以我认为这是最好的方法。

    【讨论】:

      【解决方案3】:

      这是一种非常常见的操作,pandas 有一种优雅的方法。为了避免重复 26 个求和函数的繁琐任务,我使用了字典理解。

      首先按列定义动作字典,然后使用agg 函数:

      df = pd.DataFrame({
          "Year": [1991, 1992, 1991, 1995, 1990, 1996],
          "Control": ["A1", "A1", "B1", "C1", "D1", "D1"],
          "Description": [
              "A description",
              "A description",
              "B description",
              "C description",
              "D description",
              "D description"
          ],
          "USD_Cost": [1, 100, 3, 5, 2, 1],
          "USD_Profit": [2, 300, 50, 10, 1, 1],
      })
      
      actions = {'Year': pd.Series.nunique,
                 'Description': lambda x: x.iloc[0]}
      actions.update({x: sum for x in df.columns if x.startswith('USD_')})
      
      df.groupby('Control').agg(actions).reset_index()
      

      这提供了

        Control  Year    Description  USD_Cost  USD_Profit
      0      A1     2  A description       101         302
      1      B1     1  B description         3          50
      2      C1     1  C description         5          10
      3      D1     2  D description         3           2
      

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 2021-05-24
        • 1970-01-01
        • 2018-07-19
        • 2020-09-01
        • 2017-04-29
        • 2021-12-20
        • 1970-01-01
        相关资源
        最近更新 更多