【发布时间】:2023-02-07 02:04:25
【问题描述】:
我正在尝试用 pandas 分析我的 Netflix 数据。我想总结每个用户观看特定标题的时间并打印每个配置文件的最高值。
df_clean.sample(4)
| Profile Name | Duration | time_clean |
|---|---|---|
| AAA | 0 days 00:20:00 | Harry Potter |
| AAA | 0 days 00:41:50 | The Sinner |
| BBB | 0 days 00:00:15 | Avatar |
| AAA | 0 days 00:15:00 | Harry Potter |
我只想查看每个配置文件的第一行
我尝试使用:
df_clean.groupby(['Profile Name','title_clean'])['Duration'].sum().sort_values(ascending=False).nlargest(1)
但它只返回 1 个配置文件的最大结果
| Profile Name | title_clean | |
|---|---|---|
| AAA | Harry Potter | 0 days 00:35:00 |
【问题讨论】:
-
不确定
sum()是否会按照您调用 sum 的方式削减它。 sum 已经是“求和”而不是“最高”/“最大”。对于groupby,有没有试过agg,transform。
标签: python pandas sorting group-by