【问题标题】:Datetime dtype is Object not DatetimeDatetime dtype 是 Object 而不是 Datetime
【发布时间】:2020-01-06 10:48:43
【问题描述】:

我试图通过时间戳来表示分组。首先,我必须将输入的时间(字符串)转换为日期时间。将其转换为 datetime 后,我注意到尽管给出了 pandas 添加日期的特定格式,但我不需要日期。我正在努力删除它并只保留时间对象,但我没有成功。我为删除日期所做的任何事情都会将 dtype 返回到我无法在其上执行 groupby 的对象。

示例数据:

https://miratrix.co.uk/          00:01:55
https://miratrix.co.uk/          00:02:02
https://miratrix.co.uk/          00:02:45
https://miratrix.co.uk/          00:01:22
https://miratrix.co.uk/          00:02:02
https://miratrix.co.uk/app-marketing-agency/          00:02:23
https://miratrix.co.uk/get-in-touch/          00:02:26
https://miratrix.co.uk/get-in-touch/          00:00:18
https://miratrix.co.uk/get-in-touch/          00:02:37
https://miratrix.co.uk/          00:00:31
https://miratrix.co.uk/          00:02:00
https://miratrix.co.uk/app-store-optimization-...          00:02:25
https://miratrix.co.uk/          00:03:36
https://miratrix.co.uk/app-marketing-agency/          00:02:09
https://miratrix.co.uk/get-in-touch/          00:02:14
https://?page_id=16198/          00:00:15
https://videos/channel/UCAQfRNzXGD4BQICkO1KQZUA/          00:09:07
https://miratrix.co.uk/get-in-touch/          00:01:39
https://miratrix.co.uk/app-marketing-agency/          00:01:07

到目前为止我所做的尝试

*Returned Object*
ga_organic['NEW Avg. Time on Page'] = pd.to_datetime(ga_organic['Avg. Time on Page'], format="%H:%M:%S").dt.time

*Returned Datetime but when trying to sample only time it returned an object*
ga_organic['NEW Avg. Time on Page'] = ga_organic['Avg. Time on Page'].astype('datetime64[ns]')

ga_organic['NEW Avg. Time on Page'].dt.time

我有一种关于日期时间的感觉,我不知道,这就是我遇到这个问题的原因。欢迎任何帮助或指导。

####Update####

感谢 ALollz 提供时间戳的解决方案。

ga_organic['NEW Avg. Time on Page'] = pd.to_timedelta(ga_organic['Avg. Time on Page'])

但是,当使用 GroupBy 使用此方法时,我仍然遇到相同的错误:

avg_time = ga_organic.groupby(ga_organic.index)['NEW Avg. Time on Page'].mean()

错误:“DataError:没有要聚合的数字类型”

是否有处理分组日期时间的特定功能?

【问题讨论】:

  • 如果你没有date,那么合适的dtype是timedelta64pd.to_timedelta
  • 这看起来成功了!谢谢!但是,仍然不能使用 groupby().mean()。对此有什么想法吗?
  • 您通常不会在日期时间列上使用groupby。您可能想要使用resample,因此您可以指定执行mean 操作的采样时间。请注意,resample 需要一个日期时间索引(因此您希望将''Avg. Time on Page' 列设置为索引

标签: python pandas datetime pandas-groupby python-datetime


【解决方案1】:

似乎groupby 无法将timedelta64 识别为数字类型。有几种解决方法,可以使用 numeric_only=False 或使用 total_seconds

import pandas as pd

#df = pd.read_clipboard(header=None)
#df[1] = pd.to_timedelta(df[1])

df.groupby(df.index//2)[1].mean()
#DataError: No numeric types to aggregate

# To fix pass `numeric_only=False`
df.groupby(df.index//2)[1].mean(numeric_only=False)
#0   00:01:58.500000
#1   00:02:03.500000
#2   00:02:12.500000
#3          00:01:22
#4          00:01:34
#5   00:02:12.500000
#6   00:02:52.500000
#7   00:01:14.500000
#8          00:05:23
#9          00:01:07
#Name: 1, dtype: timedelta64[ns]

使用简单的float 值和.total_seconds

df[1] = df[1].dt.total_seconds()

df.groupby(df.index//2)[1].mean()
#0    118.5
#1    123.5
#2    132.5
#3     82.0
#4     94.0
#5    132.5
#6    172.5
#7     74.5
#8    323.0
#9     67.0
#Name: 1, dtype: float64

这可以用pd.to_timedelta 指定unit='s' 转换回来

【讨论】:

  • 我得到了一个意想不到的输出。时间会发生变化,但它实际上并未对索引进行分组。 miratrix.co.uk 00:01:42 @ 00:01:42 @ 00:01:42.666666 miratrix.co.uk/get-in-touch 00:02:04.333333 miratrix.co.uk/app-store-optimization-services 00:01:30.666666 miratrix.co.uk/app-marketing-agency 00:01:42.666666 跨度>
  • hmm 重复项似乎都具有相同的值。有一种感觉,我需要做一个删除重复。想法?
  • 我只是为了说明而随机分组。不确定您真正需要在实际问题中分组的内容
  • 数据收集不正确。我可以修复它,但为了在页面估计等上给出适当的时间。我必须分组并平均时间。谢谢你的帮助。它解决了我的问题:)
猜你喜欢
  • 2018-12-26
  • 2018-07-04
  • 2012-11-19
  • 1970-01-01
  • 2022-01-19
  • 1970-01-01
  • 2014-06-04
  • 2015-06-21
  • 2016-05-29
相关资源
最近更新 更多