【问题标题】:pandas: pivot table inside multilevel dataframepandas:多级数据框中的数据透视表
【发布时间】:2016-01-15 10:11:56
【问题描述】:

我正在尝试透视表,以便在列中转换一些行值,所以从这个数据框df_behave

list 
                   date_time      field      value
 1    0 2015-05-22 05:37:59      StudentID   129
      1 2015-05-22 05:37:59      SchoolId    3
      2 2015-05-22 05:37:59      GroupeId     45
 2    3 2015-05-26 05:56:59      StudentID   129
      4 2015-05-26 05:56:59      SchoolId     65
      5 2015-05-26 05:56:59      GroupeId    13
      6 2015-05-26 05:56:59      Reference     87
 3    ......................    ......  ......

为了实现:

list 
                      date_time     StudentID   SchoolId  GroupId    Reference
     1       2015-05-22 05:37:59      129           3         45

     2      2015-05-26 05:56:59      129            65        15       87   

     3    ......................    ......  ......

使用以下代码:

def calculate():
    df_behave['value'] = df_behave['value'].astype(int)
    pi_df=pd.pivot_table(df_behave, 'value', index=['date_time'], columns='field')
    return pi_df

我试过这个:

def calculate():
    df_behave['value'] = df_behave['value'].astype(int)
    for liste, new_df in df_behave.groupby(level=0):
        pi_df=pd.pivot_table(new_df, 'value', index=['date_time'], columns='field')
        print pi_df
    return pi_df

但两人都回了我ValueError: invalid literal for long() with base 10: 'True'

【问题讨论】:

  • @Alexander 是对的,对于MultiIndex,您最好为他提到的字段reset_index 并设置并执行unstack。也许您应该过滤掉不必要的字段?

标签: python pandas pivot dataframe multi-index


【解决方案1】:

@Alexander 是对的,对于 MultiIndex,你最好 reset_index 并为他提到的字段设置并执行 unstack。也许您应该过滤掉不必要的字段?

只是一些随机样本数据:

In [308]: df
Out[308]: 
                     date_time     field  value
list index                                     
1    0     2015-05-22 05:37:59       Tom      1
     1     2015-05-22 05:37:59      Kate      2
     2     2015-05-22 05:37:59  GroupeId      3
2    3     2015-05-22 05:37:59       Tom      4
     4     2015-05-22 05:37:59      Kate      5
     5     2015-05-22 05:37:59  GroupeId      6

In [310]: df.set_index(['date_time', 'field'], append=True)\
            .reset_index('index')['value']\
            .unstack('field')
Out[310]: 
field                     GroupeId  Kate  Tom
list date_time                               
1    2015-05-22 05:37:59         3     2    1
2    2015-05-22 05:37:59         6     5    4

【讨论】:

    【解决方案2】:

    尝试重置您的索引,将其设置为list、date_time 和field,然后取消堆叠field。

    df.reset_index().set_index(['list', 'date_time', 'field']).unstack('field')
    

    由于您的 value 列似乎包含非数字数据,并且从上面的示例中它应该只包含整数,请尝试以下方法来定位您的错误数据:

    bad_rows = []
    for n in range(len(df) - 1):
        if not isinstance(df.loc[n, 'value'], int):
            bad_rows.append(n)
    

    您可能首先想尝试强制值:

    df['value'] = df['value'].astype('int')
    

    【讨论】:

    • 我检查了,没有“真”不在值中,也不在字段中!
    • MultiIndex:10191 个条目,(1.0, 0) 到 (2392.0, 12584) 数据列(共 3 列):日期 10191 个非空对象字段 10191 非空对象值 10191 非空对象数据类型:对象(3)内存使用量:395.8+ KB
    • 您的value 列是对象类型,这意味着它包含字符串等非数字数据。
    • 我认为这个 ligne df_behave['value'] = df_behave['value'].astype(int) 的问题,因为在这个分配之前我有错误DataError: No numeric types to aggregate,而在分配ValueError: invalid literal for long() with base 10: 'True' 之后,这很奇怪
    猜你喜欢
    • 2020-05-21
    • 1970-01-01
    • 1970-01-01
    • 2020-08-30
    • 2018-10-28
    • 2018-01-31
    • 1970-01-01
    • 2020-09-25
    • 1970-01-01
    相关资源
    最近更新 更多