【问题标题】:How to apply resample to a pandas Dataframe with not numerical value如何将重采样应用于没有数值的熊猫数据框
【发布时间】:2017-09-25 18:21:56
【问题描述】:

我了解到resample 不能应用非数值来自 Resampling pandas dataframe is deleting column 而且,我想将resample('30S') 应用于输入 df,如下所示:

输入_DF:

    eventTime          uuid  ts                   m_op  op.prg  op.tr     w.cycle cycle_type
0  2017-04-27 01:22:22 id1  2017-04-27 02:30:30   w     0.0     01:34:48  3       type_a                                                       
1  2017-04-27 01:23:16 id1  2017-04-27 02:31:00   w     1.0     01:33:54  3       type_a                                                      
2  2017-04-27 01:25:10 id1  2017-04-27 02:41:00   w     2.0     01:33:00  3       type_a                                                      
3  2017-04-27 01:25:32 id1  2017-04-27 02:42:45   w     3.0     01:32:00  3       type_a                                                     
4  2017-04-27 01:25:45 id1  2017-04-27 02:52:45   r     4.0     01:32:00  2       type_a                                                     

输出_DF

    eventTime          uuid  ts                   m_op  op.prg  op.tr     w.cycle cycle_type
0  2017-04-27 01:22:30 id1  2017-04-27 02:30:30   w     0.0     01:34:48  3       type_a                                                       
1  2017-04-27 01:23:00 id1  2017-04-27 02:30:30   w     0.0     01:34:48  3       type_a                                                       
2  2017-04-27 01:23:30 id1  2017-04-27 02:31:00   w     1.0     01:33:54  3       type_a                                                      
3  2017-04-27 01:24:00 id1  2017-04-27 02:31:00   w     1.0     01:33:54  3       type_a                                                      
4  2017-04-27 01:24:30 id1  2017-04-27 02:31:00   w     1.0     01:33:54  3       type_a                                                      
5  2017-04-27 01:25:00 id1  2017-04-27 02:31:00   w     1.0     01:33:54  3       type_a                                                      
6  2017-04-27 01:25:30 id1  2017-04-27 02:41:00   w     2.0     01:33:00  3       type_a                                                      
7  2017-04-27 01:26:00 id1             avg    +popular  3.5     Avg      +popular       type_a                

其中avg_of_the_values 计算相应时间的平均值,+popular 填充最流行的值或第一个 - 如果两个值具有相同的排名 - 而Avg 是通常的平均值。

我一直在应用groupBy 方法,但它不起作用。
任何建议将不胜感激。非常感谢提前.carlo

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    我实施的第一个解决方案是基于将整个问题拆分为子问题中的一个数字 - 每个非数字列都有一个 - 然后合并获得的解决方案。下面,报告一下我用来解决m_op情况的部分代码:

    smart_sub_schema_Operation_progress=["eventTime","uuid", "m_op"]    
    operation_progress_sub_smart_home_df=smart_home_df[smart_sub_schema_Operation_progress]
    operation_progress_sub_smart_home_df['m_op'] = operation_progress_sub_smart_home_df['m_op'].map({'P':0, 'W': 1, 'R': 2, 'S': 3,  'F': 4})
    operation_progress_sub_smart_home_df.eventTime = pd.to_datetime(operation_progress_sub_smart_home_df.eventTime)
    operation_progress_sub_smart_home_df.index = operation_progress_sub_smart_home_df['eventTime']
    resampled_operation_progress_sub_smart_home_df=operation_progress_sub_smart_home_df.resample('30S').reset_index()
    resampled_operation_progress_sub_smart_home_df["m_op"]=resampled_operation_progress_sub_smart_home_df["m_op"].astype(float)
    resampled_operation_progress_sub_smart_home_df.fillna(method='ffill', inplace=True)
    resampled_operation_progress_sub_smart_home_df['m_op'] = resampled_operation_progress_sub_smart_home_df['m_op'].map({0.0:'P', 1.0:'W', 2.0:'R', 3.0:'S',  4.0:'F'})
    print(resampled_operation_progress_sub_smart_home_df.to_string())
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2018-03-26
      • 2020-11-18
      • 2017-05-28
      • 2018-07-05
      • 2021-12-23
      • 2020-01-10
      • 1970-01-01
      相关资源
      最近更新 更多