【问题标题】:Is there a way to set the indices of a multi index DataFrame?有没有办法设置多索引 DataFrame 的索引?
【发布时间】:2021-12-09 05:19:00
【问题描述】:

我正在尝试创建一个多索引数据框,其中包含所有可能的索引,即使是当前不包含值的索引。我希望将这些不存在的值设置为 0。为此,我使用了以下内容:

index_levels = ['Channel', 'Duration', 'Designation', 'Manufacturing Class']

grouped_df = df.groupby(by = index_levels)[['Total Purchases', 'Sales', 'Cost']].agg('sum')

grouped_df = grouped_df.reindex(pd.MultiIndex.from_product(grouped_df.index.levels), fill_value = 0)

预期结果:

 ___________________________________________________________________________________________ 
|Chan. | Duration   | Designation|    Manufact. |Total Purchases|  Sales      |   Cost      |
|______|____________|____________|______________|_______________|_____________|_____________|
|      | Month      | Special    |    Brand     |    0          |  0.00       |    0.00     |
|      |            |            |______________|_______________|_____________|_____________|
|      |            |            |    Generic   |    1567       | 16546.07    |   16000.00  |
|Retail|____________|____________|______________|_______________|_____________|_____________|
|      | Season     |Not Special |    Brand     |     351       | 13246.00    |   15086.26  |
|      |            |            |______________|_______________|_____________|_____________|
|      |            |            |    Generic   |     0         |  0.00       |    0.00     |
|______|____________|____________|______________|_______________|_____________|_____________|

当至少一个索引级别包含一个值时,会产生此结果。但是,如果索引级别不包含任何值,则下面会产生以下结果。

___________________________________________________________________________________________ 
|Chan. | Duration   | Designation|    Manufact. |Total Purchases|  Sales      |   Cost      |
|______|____________|____________|______________|_______________|_____________|_____________|
|      | Monthly    |   Special  |    Generic   |    1567       | 16546.07    |   16000.00  |
|Retail|____________|____________|______________|_______________|_____________|_____________|
|      | Season     |Not Special |    Brand     |     351       | 13246.00    |   15086.26  |
|______|____________|____________|______________|_______________|_____________|_____________|

由于某种原因,这些值会继续被自动截断。如何修复索引,以便始终产生所需的结果,并且我始终可以可靠地使用这些索引进行计算,即使所述索引中没有值?

【问题讨论】:

    标签: python pandas multi-index


    【解决方案1】:

    您应该更改重新索引部分,因为pd.MultiIndex.from_product() 应该将原始数据帧索引作为输入(通过将grouped_df.index.levels 作为输入,您只传递 groupby 之后产生的索引)。

    这是一个可行的解决方案:

    full_idx = [df[col].dropna().unique() for col in index_levels]
    grouped_df = grouped_df.reindex(pd.MultiIndex.from_product(full_idx), fill_value = 0)
    

    如果您也对 NaN 类别感兴趣,则应在定义完整索引时删除 dropna()。

    【讨论】:

    • 不幸的是,上述解决方案只返回了与不想要的结果相同的输出。有没有办法设置索引?
    • 是的,您可以手动添加它,例如:full_idx = pd.MultiIndex.from_product([['Retail'], ['Month', 'Season'], ['Special', 'Not Special'], ['Brand', 'Generic']])。但是,请再次检查之前提出的解决方案,因为它在我制作的一些示例中有效(考虑到以前我没有在 reindex 部分重新分配 grouped_df)
    猜你喜欢
    • 2019-11-13
    • 1970-01-01
    • 2014-05-01
    • 2022-12-01
    • 1970-01-01
    • 2020-11-14
    • 1970-01-01
    • 2020-11-06
    • 1970-01-01
    相关资源
    最近更新 更多