Python pandas 从长到宽答案

【问题标题】：Python pandas pivot from long to widePython pandas 从长到宽
【发布时间】：2016-10-05 20:18:01
【问题描述】：

我的数据目前是长格式。下面是一个示例：

     Stock         Date      Time     Price     Year
       AAA   2001-01-05  15:20:09     2.380     2001
       AAA   2002-02-23  10:13:24     2.440     2002
       AAA   2002-02-27  17:17:55     2.460     2002
       BBB   2006-05-13  16:03:49     2.780     2006
       BBB   2006-10-04  10:33:10     2.800     2006

我想通过“Stock”和“Year”将其重塑为宽格式，如下所示：

     Stock   Year       Date1        Time1    Price1        Date2      Time2   Price2
       AAA   2001  2001-01-05     15:20:09     2.380
       AAA   2002  2002-02-23     10:13:24     2.440   2002-02-27   17:17:55    2.460
       BBB   2006  2006-05-13     16:03:49     2.780   2006-10-04   10:33:10    2.800

我尝试了这里发布的解决方案 Pandas long to wide reshape 并得到了这个：

df['idx'] = df.groupby(['Stock', 'Year']).cumcount()

df['date_idx'] = 'date_' + df.idx.astype(str)
df['time_idx'] = 'time_' + df.idx.astype(str)
df['price_idx'] = 'price_' + df.idx.astype(str)

date = df.pivot(index=['Stock', 'Year'], columns='date_idx', values='Date')
time = df.pivot(index=['Stock', 'Year'], columns='time_idx', values='Time')
price = df.pivot(index=['Stock', 'Year'], columns='price_idx', values='Price')

reshape = pd.concat([date, time, price], axis=1)

但最后一行给了我这个错误：

ValueError: 传递的项目数错误 15624，位置暗示 2

我的代码哪里出错了？还是有另一种更清洁的方式来进行这种重塑？

【问题讨论】：

我列举了几个例子here

标签： python pandas dataframe pivot-table reshape

【解决方案1】：

我认为你可以使用pivot_table，但需要一些aggfunc。我选择first，因为使用默认np.mean 和datetime 存在问题。

更好的示例解释是here 和docs。

解决方案1：

df['idx'] = (df.groupby(['Stock', 'Year']).cumcount() + 1).astype(str)

df1 = (df.pivot_table(index=['Stock', 'Year'], 
                      columns=['idx'], 
                      values=['Date', 'Time', 'Price'], 
                      aggfunc='first'))
df1.columns = [''.join(col) for col in df1.columns]
df1 = df1.reset_index()
print (df1)
  Stock  Year       Date1       Date2     Time1     Time2 Price1 Price2
0   AAA  2001  2001-01-05        None  15:20:09      None   2.38   None
1   AAA  2002  2002-02-23  2002-02-27  10:13:24  17:17:55   2.44   2.46
2   BBB  2006  2006-05-13  2006-10-04  16:03:49  10:33:10   2.78    2.8

然后你可以转换成floatprice列和to_datetimedate列：

cols = df1.columns[df1.columns.str.contains('Price')]
df1[cols] = df1[cols].astype(float)

cols = df1.columns[df1.columns.str.contains('Date')]
df1[cols] = df1[cols].apply(pd.to_datetime)


print (df1)
  Stock  Year      Date1      Date2     Time1     Time2  Price1  Price2
0   AAA  2001 2001-01-05        NaT  15:20:09      None    2.38     NaN
1   AAA  2002 2002-02-23 2002-02-27  10:13:24  17:17:55    2.44    2.46
2   BBB  2006 2006-05-13 2006-10-04  16:03:49  10:33:10    2.78    2.80

print (df1.dtypes)
Stock             object
Year               int64
Date1     datetime64[ns]
Date2     datetime64[ns]
Time1             object
Time2             object
Price1           float64
Price2           float64

解决方案2：

df['idx'] = df.groupby(['Stock', 'Year']).cumcount() + 1

df['date_idx'] = 'date_' + df.idx.astype(str)
df['time_idx'] = 'time_' + df.idx.astype(str)
df['price_idx'] = 'price_' + df.idx.astype(str)

date = df.pivot_table(index=['Stock', 'Year'], columns='date_idx', values='Date', aggfunc='first')
time = df.pivot_table(index=['Stock', 'Year'], columns='time_idx', values='Time', aggfunc='first')
price = df.pivot_table(index=['Stock', 'Year'], columns='price_idx', values='Price', aggfunc='first')

reshape = pd.concat([date, time, price], axis=1).reset_index()
print (reshape)
  Stock  Year      date_1      date_2    time_1    time_2  price_1  price_2
0   AAA  2001  2001-01-05        None  15:20:09      None     2.38      NaN
1   AAA  2002  2002-02-23  2002-02-27  10:13:24  17:17:55     2.44     2.46
2   BBB  2006  2006-05-13  2006-10-04  16:03:49  10:33:10     2.78     2.80

【讨论】：