【问题标题】:Python Pandas: Data DownsamplingPython Pandas:数据下采样
【发布时间】:2012-05-30 13:10:38
【问题描述】:

我的数据如下所示:

TEST
2012-05-01 00:00:00.203 OFF 0
2012-05-01 00:00:11.203 OFF 0
2012-05-01 00:00:22.203 ON 1
2012-05-01 00:00:33.203 ON 1
2012-05-01 00:00:44.203 OFF 0
TEST
2012-05-02 00:00:00.203 OFF 0
2012-05-02 00:00:11.203 OFF 0
2012-05-02 00:00:22.203 OFF 0
2012-05-02 00:00:33.203 ON 1
2012-05-02 00:00:44.203 ON 1
2012-05-02 00:00:55.203 OFF 0

最终,我希望能够将这样的数据下采样到各个日期,例如使用均值、最小值、最大值。 我无法让它为我的数据工作并收到此错误:

TypeError: unhashable type: 'list'

也许它与数据框中的日期格式有关,因为索引行如下所示:

[datetime.datetime(2012, 5, 1, 0, 0, 0, 203000)]   OFF  0

谁能帮忙。 到目前为止我的代码是这样的:

import time
import dateutil.parser
from pandas import *
from pandas.core.datetools import *



t0 = time.clock()

filename = "testdata.dat"

index = []
data = []

with open(filename) as f:
    for line in f:
        if not line.startswith('TEST'):
            line_content =  line.split(' ')

            mydatetime =  dateutil.parser.parse(line_content[0] +  " " + line_content[1])

            del line_content[0] # delete the date
            del line_content[0] # delete the time so that only values remain

            index_row = [mydatetime]
            data_row = []
            for item in line_content:
                data_row.append(item)

            index.append(index_row)
            data.append(data_row)


df = DataFrame(data, index = index)
print df.head()
print df.tail()

print
date_from =  index[0] # first datetime entry in data frame
print date_from
date_to =  index[len(index)-1] #last datetime entry in date frame
print date_to

print date_to[0] - date_from[0]
dayly= DateRange(date_from[0], date_to[0], offset=datetools.DateOffset())
print dayly

grouped = df.groupby(dayly.asof)
#print grouped.mean()
#df2 = df.groupby(daily.asof).agg({'2':np_mean})


time2 = time.clock() - t0
print time2

【问题讨论】:

  • 尊敬的用户1412286,请提供错误输出以获得有效帮助

标签: python pandas downsampling


【解决方案1】:

您最好将所有日期时间插值留给pandas,并用干净的输入流来提供它。然后,您可以使用 read_fwf 分隔字段(对于固定宽度的格式化行)。例如:

import pandas
import StringIO

buf = StringIO.StringIO()
buf.write(''.join(line
    for line in open('f.txt')
    if not line.startswith('TEST')))
buf.seek(0)

df = pandas.read_fwf(buf, [(0, 24), (24, 27), (27, 30)],
        index_col=0, names=['switch', 'value'])
print df

输出:

                        switch  value
2012-05-01 00:00:00.203    OFF      0
2012-05-01 00:00:11.203    OFF      0
2012-05-01 00:00:22.203     ON      1
2012-05-01 00:00:33.203     ON      1
2012-05-01 00:00:44.203    OFF      0
2012-05-02 00:00:00.203    OFF      0
2012-05-02 00:00:11.203    OFF      0
2012-05-02 00:00:22.203    OFF      0
2012-05-02 00:00:33.203     ON      1
2012-05-02 00:00:44.203     ON      1
2012-05-02 00:00:55.203    OFF      0

【讨论】:

  • 数据不一定总是包含相同数量的列,如果可能的话,我想避免每次读取新文件时都必须手动调整代码。
  • 你不需要。只需使用read_table 或read_csv 或read_fwf,具体取决于您期望的格式。如果您收到的文件没有格式,那么我几乎看不到自动解析的方法!
  • 它们确实有一种格式,即时间戳始终存在,但数据列的数量可能会有所不同。到目前为止,我还无法使用 read_csv 正确读取时间戳,可能是因为日期和时间之间存在空格,因此与其他列没有区别。或者让我更具体一点:我已经能够通过从每一行创建一个列表然后将其附加到另一个列表来正确读取时间戳,但我还没有设法将时间戳作为数据帧的索引。
  • 再次,read_table 可以使用,例如:pandas.read_table(buf, sep=' ', index_col=[0,1], header=None) 将创建一个包含多个列的表,multiindex 由 2 级组成:第一级年-月-日,第二级平时间。如果您愿意,您可以将多索引合并到普通索引中(例如:df.index = ['%s %s' % (a, b) for a, b in zip(df.index.get_level_values(0), df.index.get_level_values(1))]。)
  • 太棒了。如果您可以批准答案,则可以将其关闭。
【解决方案2】:

我对@9​​87654321@ 没有任何经验,但从我的代码中可以看出,

df = DataFrame(data, index = index)

和错误,似乎index 不应该是像 python 列表这样的可变对象。也许这会起作用:

df = DataFrame(data, index = tuple(index))

此外,您的 index_row 和 data_row 本身就是列表似乎并不明显,而您将它们附加到 index 和 data 列表中。

【讨论】:

  • 不,这不起作用。最初,所有列都在一个列表“数据”中,然后将其转换为数据框。然后,日期和时间正确显示。但要使下采样工作,它们需要在索引中。
猜你喜欢
  • 2018-02-01
  • 2017-01-27
  • 1970-01-01
  • 2013-03-27
  • 1970-01-01
  • 2015-12-27
  • 2015-02-20
  • 2019-04-15
  • 2023-03-13
相关资源
最近更新 更多