【问题标题】:Pandas Create a column with the a sum of a nested dataframe column熊猫用嵌套数据框列的总和创建一列
【发布时间】:2020-04-19 21:00:31
【问题描述】:

如何在不丢失任何其他列和使用 pandas 的嵌套数据的情况下将新列添加到包含嵌套数据帧中的值总和的数据帧?

具体来说,我想创建一个新列total_cost,其中包含一行的所有嵌套数据框的总和。

我设法使用一系列groupby 和apply 创建了以下数据框:

  user_id   description                                       unit_summary
0     111  xxx  [{'total_period_cost': 100, 'unit_id': 'xxx', ...
1     222  xxx  [{'total_period_cost': 100, 'unit_id': 'yyy', ...

我正在尝试添加列total_cost,它是每个嵌套数据框的total_period_cost 的总和(按user_id 分组)。如何实现以下数据框?

  user_id   description   total_cost                          unit_summary
0     111  xxx  300  [{'total_period_cost': 100, 'unit_id': 'xxx', ...
1     222  xxx  100  [{'total_period_cost': 100, 'unit_id': 'yyy', ...

我的代码:

import pandas as pd

series = [{
    "user_id":"111", 
    "description": "xxx",
    "unit_summary":[
        {
        "total_period_cost":100,
        "unit_id":"xxx",
        "cost_per_unit":50,
        "total_period_usage":2
        },
        {
        "total_period_cost":200,
        "unit_id":"yyy",
        "cost_per_unit":25,
        "total_period_usage": 8
        }
    ]
},
{
    "user_id":"222",
    "description": "xxx",
    "unit_summary":[
        {
            "total_period_cost":100,
            "unit_id":"yyy",
            "cost_per_unit":25,
            "total_period_usage": 4
        }
    ]
}]

df = pd.DataFrame(series)

print(df)
print(df.to_dict(orient='records'))

下面是我用来实现series JSON对象的groupby..apply代码示例:

import pandas as pd

series = [
    {"user_id":"111", "unit_id":"xxx","cost_per_unit":50, "total_period_usage": 1},
    {"user_id":"111", "unit_id":"xxx","cost_per_unit":50, "total_period_usage": 1},
    {"user_id":"111", "unit_id":"yyy","cost_per_unit":25, "total_period_usage": 8},
    {"user_id":"222", "unit_id":"yyy","cost_per_unit":25, "total_period_usage": 3},
    {"user_id":"222", "unit_id":"yyy","cost_per_unit":25, "total_period_usage": 1}
]

df = pd.DataFrame(series)

sumc = (
    df.groupby(['user_id', 'unit_id', 'cost_per_unit'], as_index=False)
        .agg({'total_period_usage': 'sum'})
)

sumc['total_period_cost'] = sumc.total_period_usage * sumc.cost_per_unit

sumc = (
    sumc.groupby(['user_id'])
        .apply(lambda x: x[['total_period_cost', 'unit_id', 'cost_per_unit', 'total_period_usage']].to_dict('r'))
        .reset_index()
)

sumc = sumc.rename(columns={0:'unit_summary'})

sumc['description'] = 'xxx'

print(sumc)
print(sumc.to_dict(orient='records'))

通过从 anky_91 的答案中添加以下内容来解决它:

def myf(x):
    return pd.DataFrame(x).loc[:,'total_period_cost'].sum()
# Sum all server sumbscriptions total_period_cost
sumc['total_period_cost'] = sumc['unit_summary'].apply(myf)

【问题讨论】:

  • 您不应该编写这样的代码,永远不需要将groupby..apply 的中间结果存储为嵌套数据框,它只会让您的生活更加艰难。 When you want to add a new (summary) column but also keep the existing rows and columns, use .transform()。如果您发布您的原始groupby..apply 代码,我们可以正确实现。
  • 我发布了我原来的 groupby..apply,现在我使用 .apply() 和现有 group by 的函数得到了预期的结果。我阅读了转换,但我还不知道如何使用它。
  • @barracuda 您是否要将单元摘要用于其他目的?如果不是,我在答案末尾添加了另一种方法
  • @anky_91 Method1 完美地创建了我正在寻找的最终数据集。将数据框转换为 JSON 对象时,我使用 unit_summary 在最终 JSON 对象中创建嵌套数组。

标签: python pandas dataframe


【解决方案1】:

您可以将unit_summary 列中的每一行作为数据框读取并求和所需的列:

方法一: apply

def myf(x):
    return pd.DataFrame(x).loc[:,'total_period_cost'].sum()
df['total_cost'] = df['unit_summary'].apply(myf)

print(df)

方法二: 类似地通过列表理解:

df['total_cost'] = [pd.DataFrame(i)['total_period_cost'].sum() 
                            for i in df['unit_summary'].tolist()]

方法3:使用explode:

m = df['unit_summary'].explode()
df['total_cost'] = pd.DataFrame(m.tolist(),index=m.index)['total_period_cost'].sum(level=0)

  user_id description                                       unit_summary  \
0     111         xxx  [{'total_period_cost': 100, 'unit_id': 'xxx', ...   
1     222         xxx  [{'total_period_cost': 100, 'unit_id': 'yyy', ...   

   total_cost  
0         300  
1         100 

除了上述之外,从您的原始数据框开始,我们还可以执行以下操作来实现所需的输出,但是这不会为您提供带有 dicts ('unit_summary`) 的系列:

(df.assign(total_cost=df['cost_per_unit']*df['total_period_usage'])
  .groupby(['user_id'],as_index=False)['total_cost'].sum().assign(description='xxxx'))

  user_id  total_cost description
0     111         300        xxxx
1     222         100        xxxx

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-04-06
    • 2023-03-27
    • 2018-03-07
    • 2019-01-11
    • 2020-07-26
    • 2017-01-01
    • 1970-01-01
    相关资源
    最近更新 更多