【问题标题】:Sort out the data in different group and calculate the start time & end in Pandas对不同组中的数据进行排序并计算 Pandas 中的开始时间和结束时间
【发布时间】:2017-06-16 20:25:47
【问题描述】:

我做了一个cpu功耗的性能测试,得到了一组csv格式的数据。从数据中,有 5 个不同的事件,我想整理出每个事件并计算每个事件的开始时间和结束时间。我尝试在 Python 中使用 Pandas 进行数据分析,但是,我仍然不知道该怎么做。以下是我到目前为止编写的非常基本的代码。

import pandas as pd
from pandas import DataFrame
import os, sys

df = pd.read_csv('new.csv')

col_Time = df[df.columns[0]]
col_Data = df[df.columns[1]]

## example_time = [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15]
## example_data_in_watts=[11.2, 10.3, 10.1, 21.2, 20.3, 22.1, 12.3, 10.7, 
##                        11.2, 23.6, 24.3, 25.1, 10.2, 11.3, 10.5]

## As above, each element in example_data_in_watts corresponds to element in 
## example_time. From this data, there're 2 events happened when the watts
## are ~21w and ~24w. My desired output will be to calculate the start & end 
## time for 21w & 24w, which are 3(sec) and 3(sec). 

正如你在上面看到的,我只分配了两个变量来代表2个不同的列:一个是测试时间(单位:秒),另一个是测试数据(单位:瓦特)。我能想到的一种方法可能是使用 k-means 方法来整理事件。但即使我这样做,我也不确定是否可以从那里获得开始时间和结束时间?

如果有人知道如何整理事件并计算开始时间和结束时间,请告诉我。非常感谢!!

【问题讨论】:

  • 如果您可以包含示例输入和所需的输出,对您的帮助会容易得多。
  • 谢谢! :) 我刚刚在上面的代码中包含了示例输入和所需的输出。

标签: python csv pandas


【解决方案1】:

从您的示例来看,这似乎是最简单的解决方案:

watts=[10,10,10,21,21,21,10,10,10,23,23,23,10,10,10]

result = {k : watts[:i+1].count(k) for i, k, in enumerate(watts) if k != 10}

编辑

如果您有浮动数据,根据您的示例,您可以执行以下操作:

watts=[10.2, 10.3, 10.1, 21.2, 21.3, 21.1, 10.3, 10.7, 10.2, 23.6, 23.3, 23.1, 10.2, 10.3, 10.5]
watts = map(int,watts)

result = {k : watts[:i+1].count(k) for i, k, in enumerate(watts) if k != 10}

编辑编辑

考虑到变化是 2.5,我认为这可以解决问题:

watts=[11.2, 10.3, 10.1, 21.2, 20.3, 22.1, 12.3, 10.7, 11.2, 23.6, 24.3, 25.1, 10.2, 11.3, 10.5]

watts = map(lambda x: x + x % 5, map(lambda x: x - x % 2.5, map(int, watts)))

result = {k : watts[:i+1].count(k) for i, k, in enumerate(watts) if k != 10}

【讨论】:

  • 谢谢!但是数据集实际上要复杂得多,因为功率差异很大。所以它就像从 [10.1, 10.3, 10.2] 跳到 [21.1, 21.4, 21.5] 等等。
  • 谢谢! :) 但是很抱歉,我仍然没有把这个例子说清楚。请检查我上面编辑的代码。更具体地说,我想知道有什么方法可以整理出这两个事件(组)并计算每个事件花费了多少秒?因为原始数据量巨大,而power值实际上有一些小的波动。
  • 我这边的另一种解决方法,这肯定不是正确的方法,但它可能会起作用:)
  • 它仍然无法正常工作......代码给了我一个错误“TypeError:'map' object is not subscriptable”。我认为这可能是由于我使用的是 Python 3。此外,我认为我仍然需要一种更“正确的方法”来分析数据。
猜你喜欢
  • 2021-04-10
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2020-07-01
  • 1970-01-01
相关资源
最近更新 更多