【问题标题】:Calculating total age of a values based on time column根据时间列计算值的总年龄
【发布时间】:2020-06-25 05:33:10
【问题描述】:

我有一个带有列的 DataFrame:-

Time                   channel-id     user-id   call-id   call-duration
2020-06-25 00:06:22    abc            dg33      3532      30
2020-06-25 00:06:34    dfd            sd24      2342      35
2020-06-25 00:07:23    abs            gf22      5467      40
2020-06-25 00:07:44    abc            sd33      9233      42
2020-06-25 00:07:23    dfd            sd24      4938      2
2020-06-25 00:08:44    abs            hg34      2933      55
2020-06-25 00:08:43    abc            lk33      2933      43
2020-06-25 00:08:00    dfd            sd11      4532      56
2020-06-25 00:09:11    abc            lf33      2283      76
2020-06-25 00:09:12    abc            df43      4466      23
2020-06-25 00:09:55    abc            cv45      8888      12

我想计算频道的生命周期,例如频道abc 开始于2020-06-25 00:06:22,但在频道abc 中,最后一个用户加入2020-06-25 00:09:55。

我想列出所有频道以及每个频道的频道生命周期。

channel-id  channel-life-duration
abc         5 minutes
xyz         300 minutes

我提到 5 分钟和 300 分钟只是为了说明格式。

另外,如果可能的话,我还想计算总的和唯一的用户 ID 和呼叫 ID。

channel-id  channel-life-duration   Total-user-id  Tot-unique-user-id  Total-call-id  tot-unique-call-id
abc         5 minutes                10             6                   11             10       
xyz         300 minutes              11             7                   12             9        

我想把它扩展到数百万行,所以我怎样才能使计算速度更快。

【问题讨论】:

  • @jezrael 你能帮忙吗!
  • 你能解释一下你是如何在abc 的持续时间内得到 5 分钟的吗?
  • @ShubhamSharma 我没明白!我是说我想以同样的方式。我刚刚提到的 5 分钟给出了预期输出的想法

标签: python pandas numpy dataframe lambda


【解决方案1】:

在channel-id上使用DataFrame.groupby,然后根据要求使用.agg聚合分组数据帧:

df1 = (
    df.groupby('channel-id').agg(
        first=('Time', 'first'), last=('Time', 'last'),
        user_id=('user-id', 'count'), unique_user_id=('user-id', 'nunique'),
        call_id=('call-id', 'count'), unique_call_id=('call-id', 'nunique'),
        call_duration=('call-duration', 'last'))
    .add_prefix('total_')
    .rename(columns={'total_last': 'channel_life_duration'})
)

# Calculate the lifespan of the channel
df1['channel_life_duration'] = (
    df1['channel_life_duration'].sub(df1.pop('total_first'))
    .add(pd.to_timedelta(df1.pop('total_call_duration'), unit='s'))
    .div(np.timedelta64(1, 'm'))
)

结果:

# print(df1)

            channel_life_duration  total_user_id  total_unique_user_id  total_call_id  total_unique_call_id
channel-id                                                                                                 
abc                      3.750000              6                     6              6                     6
abs                      2.266667              2                     2              2                     2
dfd                      2.366667              3                     2              3                     3

【讨论】:

  • 它给了我 typeerror :- TypeError: aggregate() missing 1 required positional argument: 'arg' @shubham
  • 我会升级试试
  • 好的,非常感谢!一个问题?是 scalabel,我的意思是,如果我将它用于数十亿条记录,那么有什么建议可以让它快速或良好吗?
  • 我想应该没问题,因为我们在进行所有计算时都使用了矢量化通用函数。
  • @abhi 当然。我已经编辑了答案,现在可以为您提供新的所需结果。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2019-12-12
  • 2011-05-26
  • 1970-01-01
  • 1970-01-01
  • 2019-10-19
  • 2013-10-31
相关资源
最近更新 更多