【问题标题】:Why is my data not recognized as time series?为什么我的数据不被识别为时间序列?
【发布时间】:2017-02-21 12:01:24
【问题描述】:

我有一个人 (cal2) 的每日 (day) 卡路里摄入量数据,这些数据来自 Stata dta 文件。

我运行下面的代码:

import pandas as pd
import numpy as np
import matplotlib.pylab as plt
from pandas import read_csv
from matplotlib.pylab import rcParams

d = pd.read_stata('time_series_calories.dta', preserve_dtypes=True, 
                  index = 'day', convert_dates=True)

print(d.dtypes)
print(d.shape)
print(d.index)
print(d.head)

plt.plot(d)

这是数据的样子:

0   2002-01-10  3668.433350
1   2002-01-11  3652.249756
2   2002-01-12  3647.866211
3   2002-01-13  3646.684326
4   2002-01-14  3661.941406
5   2002-01-15  3656.951660

印刷品显示以下内容:

day     datetime64[ns]
cal2           float32
dtype: object

(251, 2)

Int64Index([  0,   1,   2,   3,   4,   5,   6,   7,   8,   9,
            ...
            241, 242, 243, 244, 245, 246, 247, 248, 249, 250],
           dtype='int64', length=251)

这就是问题所在 - 数据应标识为 dtype='datatime64[ns]'。

但是,显然不是。为什么不呢?

【问题讨论】:

  • 除非我们看到数据是什么样子,否则无法提供帮助。只需要几行。
  • 当然!已编辑。谢谢!
  • 这就是“csv”的样子吗?抱歉,'dta'?
  • 在打印语句中,它打印“datetime64[ns]”,即您要查找的内容....?
  • 不,我的来源是 .dta - 不幸的是不是 .csv。但是print.(d.head) 返回上面显示的结构。

标签: python-3.x pandas import time-series stata


【解决方案1】:

提供的代码、数据和显示的类型之间存在差异。 这是因为不管cal2 的类型如何,index = 'day' 参数 在pd.read_stata() 中应该始终呈现day 索引,尽管不是作为 想要的类型。

话虽如此,问题可以重现如下。

首先,在Stata中创建数据集:

clear
input double day float cal2
15350  3668.433
15351   3652.25
15352  3647.866
15353  3646.684
15354 3661.9414
15355  3656.952
end
format %td day

save time_series_calories

describe

Contains data from time_series_calories.dta
  obs:             6                          
 vars:             2                          
 size:            72                          
----------------------------------------------------------------------------------------------------
              storage   display    value
variable name   type    format     label      variable label
----------------------------------------------------------------------------------------------------
day             double  %td                   
cal2            float   %9.0g                 
----------------------------------------------------------------------------------------------------
Sorted by: 

其次,在 Pandas 中加载数据:

import pandas as pd
d = pd.read_stata('time_series_calories.dta', preserve_dtypes=True, convert_dates=True)

print(d.head)
         day         cal2
0 2002-01-10  3668.433350
1 2002-01-11  3652.249756
2 2002-01-12  3647.866211
3 2002-01-13  3646.684326
4 2002-01-14  3661.941406
5 2002-01-15  3656.951660

print(d.dtypes)
day     datetime64[ns]
cal2           float32
dtype: object

print(d.shape)
(6, 2)

print(d.index)
Int64Index([0, 1, 2, 3, 4, 5], dtype='int64')

为了根据需要更改索引,您可以使用pd.set_index():

d = d.set_index('day')

print(d.head)

                   cal2
day                    
2002-01-10  3668.433350
2002-01-11  3652.249756
2002-01-12  3647.866211
2002-01-13  3646.684326
2002-01-14  3661.941406
2002-01-15  3656.951660

print(d.index)
DatetimeIndex(['2002-01-10', '2002-01-11', '2002-01-12', '2002-01-13',
               '2002-01-14', '2002-01-15'],
              dtype='datetime64[ns]', name='day', freq=None)

如果day是Stata数据集中的一个字符串,那么你可以这样做:

d['day'] = pd.to_datetime(d.day)
d = d.set_index('day') 

【讨论】:

  • 瑞秋,我的回答直接解决了你的问题。所以请考虑投票并接受它。谢谢。
猜你喜欢
  • 1970-01-01
  • 2019-10-23
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2015-04-28
  • 1970-01-01
  • 1970-01-01
  • 2021-02-22
相关资源
最近更新 更多