【问题标题】:Issue with Pandas resampling熊猫重采样问题
【发布时间】:2015-11-21 12:37:36
【问题描述】:

我有一个字典列表如下:

>>>L=[
   {
   "timeline": "2014-10", 
   "total_prescriptions": 17
   }, 
   {
   "timeline": "2014-11", 
   "total_prescriptions": 14
   }, 
   {
   "timeline": "2014-12", 
   "total_prescriptions": 8
  },
  {
  "timeline": "2015-1", 
  "total_prescriptions": 4
  }, 
  {
  "timeline": "2015-3", 
  "total_prescriptions": 10
  }, 
  {
  "timeline": "2015-4", 
  "total_prescriptions": 3
  } 
  ]

我需要做的是填补缺失的月份,在这种情况下是 2015 年 2 月,总处方为零。我使用 Pandas 如下:

>>> df = pd.DataFrame(L)
>>> df.index=pd.to_datetime(df.timeline,format='%Y-%m')
>>> df
           timeline  total_prescriptions
timeline
2014-10-01  2014-10                  17
2014-11-01  2014-11                  14
2014-12-01  2014-12                   8
2015-01-01  2015-1                    4
2015-03-01  2015-3                   10
2015-04-01  2015-4                    3

>>> df = df.resample('MS').fillna(0)
>>> df
            total_prescriptions
timeline
2014-10-01                   17
2014-11-01                   14
2014-12-01                    8
2015-01-01                    4
2015-02-01                    0
2015-03-01                   10
2015-04-01                    3

到目前为止一切顺利..正是我想要的..现在我需要将此数据框转换回字典列表..我就是这样做的:

>>> response = df.T.to_dict().values()
>>> response
[{'total_prescriptions': 0.0}, 
 {'total_prescriptions': 17.0},     
 {'total_prescriptions': 10.0}, 
 {'total_prescriptions': 14.0}, 
 {'total_prescriptions': 4.0}, 
 {'total_prescriptions': 8.0}, 
 {'total_prescriptions': 3.0}]

顺序丢失,时间线丢失,total_prescriptions 变成 int 的十进制值。出了什么问题?

【问题讨论】:

  • 那么 decimal 值是因为由于重采样引入的NaN 行,您的 dtype 将转换为浮点数,您可以使用以下方法转换回来:df = df.resample('MS').fillna(0).astype(np.int32) as对于排序丢失,这是由于当您调用 values 时字典不能保证顺序,您必须对键进行排序并从排序后的键构建字典

标签: python list dictionary pandas


【解决方案1】:

首先,由于重采样,到 十进制 的转换实际上是 float dtype,因为这将为缺失值引入 NaN 值,您可以使用 astype 修复此问题,然后您可以恢复您的“时间线”列丢失了,因为它无法弄清楚如何重新采样str,因此我们可以将strftime 应用于索引:

In [80]:
df = df.resample('MS').fillna(0).astype(np.int32)
df['timeline'] = df.index.to_series().apply(lambda x: dt.datetime.strftime(x, '%Y-%m'))
df

Out[80]:
            total_prescriptions timeline
timeline                                
2014-10-01                   17  2014-10
2014-11-01                   14  2014-11
2014-12-01                    8  2014-12
2015-01-01                    4  2015-01
2015-02-01                    0  2015-02
2015-03-01                   10  2015-03
2015-04-01                    3  2015-04

现在我们需要对 dict 键进行排序,因为调用 values 会丢失排序顺序,我们可以执行列表推导以恢复原始形式:

In [84]:
d = df.T.to_dict()
[d[key[0]] for key in sorted(d.items())]

Out[84]:
[{'timeline': '2014-10', 'total_prescriptions': 17},
 {'timeline': '2014-11', 'total_prescriptions': 14},
 {'timeline': '2014-12', 'total_prescriptions': 8},
 {'timeline': '2015-01', 'total_prescriptions': 4},
 {'timeline': '2015-02', 'total_prescriptions': 0},
 {'timeline': '2015-03', 'total_prescriptions': 10},
 {'timeline': '2015-04', 'total_prescriptions': 3}]

【讨论】:

  • @Edchum..这有帮助..但是..时间轴值仍然缺失..目标是拥有一个像原始输入一样的数据结构..
猜你喜欢
  • 1970-01-01
  • 2016-11-25
  • 1970-01-01
  • 1970-01-01
  • 2021-06-01
  • 2013-06-04
  • 2017-10-09
  • 2022-06-10
  • 2022-01-24
相关资源
最近更新 更多