【问题标题】:Sum of values based on a condition in pandas基于熊猫条件的值总和
【发布时间】:2021-03-01 16:12:54
【问题描述】:
    ID        change         SX      Supresult
0  UNITY        NaN           0        NaN
1  UNITY    -0.009434        100     -0.015283 (P1)
2  UNITY     0.003463         0        NaN
3  TRINITY   0.008628        100     -0.043363 
4  TRINITY  -0.027374        100      0.008423 (P2)
5  TRINITY  -0.011002         0        NaN
6  TRINITY  -0.004987        100       NaN
7  TRINITY   0.007566         0        NaN

如果“SX”等于 100,我使用以下程序创建一个新列“Supresult”。新列存储 NEXT 三个“更改”值的总和。例如,在索引 1 中,supresult 是索引 2,3 和 4 变化的总和。

df['Supresult'] = df[df.SX == 100].index.to_series().apply(lambda x: df.change.shift(-1).iloc[x: x + 3].sum())

但是,我面临两个需要帮助的问题:

(P1):我希望总和是特定于“ID”的。例如,索引 1 中的结果继续前进,并从 UNITY 中获取一个值,从 TRINITY 中获取两个值的总和。只要它在同一个“ID”内,就应该进行总和。我试图在我的代码末尾添加.groupby('ID'),但它给出了一个关键错误。

(P2):由于索引 3 已经给出了接下来三个变化的总和,所以索引 4 不应该继续计算接下来三天的总和。下一个总和仅应在上一个计算周期完成后进行,即索引 6 及以后。

预期结果:

    ID        change         SX      Supresult
0  UNITY        NaN           0        NaN
1  UNITY    -0.009434        100       NaN
2  UNITY     0.003463         0        NaN
3  TRINITY   0.008628        100     -0.043363 
4  TRINITY  -0.027374        100       NaN
5  TRINITY  -0.011002         0        NaN
6  TRINITY  -0.004987        100       NaN
7  TRINITY   0.007566         0        NaN

将不胜感激,谢谢!

【问题讨论】:

    标签: python pandas pandas-groupby


    【解决方案1】:

    鉴于您的复杂要求,我认为循环是合适的:

    # If your data frame is not indexed sequentially, this will make it so.
    # The algorithm needs the frame to be indexed 0, 1, 2, ...
    df.reset_index(inplace=True)
    
    # Every row starts off in "unconsumed" state
    consumed = np.repeat(0, len(df))
    result = np.repeat(np.nan, len(df))
    
    for i, sx in df['SX'].iteritems():
        # The next three rows
        slc = slice(i+1, i+4)
    
        # A row is considered a match if:
        #   * It has SX == 100
        #   * The next three rows have the same ID
        #   * The next three rows are not involved in a previous summation
        match = (
            (sx == 100) and
            (df.loc[slc, 'ID'].nunique() == 1) and
            (consumed[i] == 0)
        )
        if match:
            consumed[slc] = 1
            result[i] = df.loc[slc, 'Supresult'].sum()
    
    df['Supresult'] = result
    

    【讨论】:

    • 代码在你端工作吗?它为我吐出“KeyError:'Supresult'”
    猜你喜欢
    • 1970-01-01
    • 2017-05-28
    • 1970-01-01
    • 2018-07-15
    • 2018-04-17
    • 2018-11-19
    • 2021-11-05
    • 2017-07-03
    • 1970-01-01
    相关资源
    最近更新 更多