【问题标题】:Convert pandas dataframe of lists to dict of dataframes将列表的熊猫数据框转换为数据框的字典
【发布时间】:2017-10-02 05:52:03
【问题描述】:

我有一个数据框(带有 DateTime 索引),其中一些列包含列表,每个列表有 6 个元素。

In: dframe.head()
Out: 
                           A                                        B  \
timestamp                                                                
2017-05-01 00:32:25        30  [-3512, 375, -1025, -358, -1296, -4019]   
2017-05-01 00:32:55        30  [-3519, 372, -1026, -361, -1302, -4020]   
2017-05-01 00:33:25        30  [-3514, 371, -1026, -360, -1297, -4018]   
2017-05-01 00:33:55        30  [-3517, 377, -1030, -363, -1293, -4027]   
2017-05-01 00:34:25        30  [-3515, 372, -1033, -361, -1299, -4025]   
                                                      C           D
timestamp                                                             
2017-05-01 00:32:25  [1104, 1643, 625, 1374, 5414, 2066]      49.93   
2017-05-01 00:32:55  [1106, 1643, 622, 1385, 5441, 2074]      49.94   
2017-05-01 00:33:25  [1105, 1643, 623, 1373, 5445, 2074]      49.91   
2017-05-01 00:33:55  [1105, 1646, 620, 1384, 5438, 2076]      49.91   
2017-05-01 00:34:25  [1104, 1645, 613, 1374, 5431, 2082]      49.94   

我有一本字典 dict_of_dfs,我想用 6 个数据框填充它,

dict_of_dfs = {1: df1, 2:df2, 3:df3, 4:df4, 5:df5, 6:df6}

其中 ith 数据框包含每个列表中的 ith 项,因此 dict 中的第一个数据框将是:

In:df1
Out: 
                            A          B      C        D
    timestamp                                                                
    2017-05-01 00:32:25        30  -3512   1104    49.93
    2017-05-01 00:32:55        30  -3519   1106    49.94
    2017-05-01 00:33:25        30  -3514   1105    49.91
    2017-05-01 00:33:55        30  -3517   1105    49.91
    2017-05-01 00:34:25        30  -3515   1104    49.94

等等。 实际的数据框有比这更多的列和数千行。 进行转换的最简单、最 Python 的方法是什么?

【问题讨论】:

    标签: python list pandas dictionary dataframe


    【解决方案1】:

    您可以将字典理解与assign 一起使用,对于lists 的选择值,请使用str[0]、str[1]:

    N = 6
    dfs = {i:df.assign(B=df['B'].str[i-1], C=df['C'].str[i-1]) for i in range(1,N + 1)}
    
    print(dfs[1])
                 timestamp   A     B     C      D
    0  2017-05-01 00:32:25  30 -3512  1104  49.93
    1  2017-05-01 00:32:55  30 -3519  1106  49.94
    2  2017-05-01 00:33:25  30 -3514  1105  49.91
    3  2017-05-01 00:33:55  30 -3517  1105  49.91
    4  2017-05-01 00:34:25  30 -3515  1104  49.94
    

    另一种解决方案:

    dfs = {i:df.apply(lambda x: x.str[i-1] if type(x.iat[0]) == list else x) for i in range(1,7)}
    
    print(dfs[1])
                 timestamp   A     B     C      D
    0  2017-05-01 00:32:25  30 -3512  1104  49.93
    1  2017-05-01 00:32:55  30 -3519  1106  49.94
    2  2017-05-01 00:33:25  30 -3514  1105  49.91
    3  2017-05-01 00:33:55  30 -3517  1105  49.91
    4  2017-05-01 00:34:25  30 -3515  1104  49.94
    

    时间安排:

    df = pd.concat([df]*10000).reset_index(drop=True)
    
    In [185]: %timeit {i:df.assign(B=df['B'].str[i-1], C=df['C'].str[i-1]) for i in range(1,N+1)}
    1 loop, best of 3: 420 ms per loop
    
    In [186]: %timeit {i:df.apply(lambda x: x.str[i-1] if type(x.iat[0]) == list else x) for i in range(1,7)}
    1 loop, best of 3: 447 ms per loop
    
    In [187]: %timeit {(i+1):df.applymap(lambda x: x[i] if type(x) == list else x) for i in range(6)}
    1 loop, best of 3: 881 ms per loop
    

    【讨论】:

    • 我添加了另一个解决方案,没有指定列名,请检查。
    • 看起来不错,与上面@Allen 的解决方案相同。它在控制台中也适用于我,但是当包含在我的脚本中时,我得到 NameError: ("free variable 'type' referenced before assignment in enclosing scope", 'occurred at index duration') 知道为什么吗??
    • 我认为有一些变量被称为type,这也是python中的代码字。所以尝试更改变量名。 Ans 解决方案有点改变,因为更快;)
    • 我同意这就是它的样子,但我的脚本中没有名为 type 的变量
    • 是的,有错误。万岁!!!!!!这种类型的错误真的很糟糕:(
    【解决方案2】:

    设置

    df = pd.DataFrame({'A': {'2017-05-01 00:32:25': 30,
      '2017-05-01 00:32:55': 30,
      '2017-05-01 00:33:25': 30,
      '2017-05-01 00:33:55': 30,
      '2017-05-01 00:34:25': 30},
     'B': {'2017-05-01 00:32:25': [-3512, 375, -1025, -358, -1296, -4019],
      '2017-05-01 00:32:55': [-3519, 372, -1026, -361, -1302, -4020],
      '2017-05-01 00:33:25': [-3514, 371, -1026, -360, -1297, -4018],
      '2017-05-01 00:33:55': [-3517, 377, -1030, -363, -1293, -4027],
      '2017-05-01 00:34:25': [-3515, 372, -1033, -361, -1299, -4025]},
     'C': {'2017-05-01 00:32:25': [1104, 1643, 625, 1374, 5414, 2066],
      '2017-05-01 00:32:55': [1106, 1643, 622, 1385, 5441, 2074],
      '2017-05-01 00:33:25': [1105, 1643, 623, 1373, 5445, 2074],
      '2017-05-01 00:33:55': [1105, 1646, 620, 1384, 5438, 2076],
      '2017-05-01 00:34:25': [1104, 1645, 613, 1374, 5431, 2082]},
     'D': {'2017-05-01 00:32:25': 49.93,
      '2017-05-01 00:32:55': 49.94,
      '2017-05-01 00:33:25': 49.1,
      '2017-05-01 00:33:55': 49.91,
      '2017-05-01 00:34:25': 49.94}})
    

    解决方案

    使用字典理解构造 df 字典。 sub df 是使用 applymap 函数生成的。它可以转换包含 6 个元素的列表的所有列:

    dict_of_dfs = {(i+1):df.applymap(lambda x: x[i] if type(x) == list else x) for i in range(6)}
    
    print(dict_of_dfs[1])
                          A     B     C      D
    2017-05-01 00:32:25  30 -3512  1104  49.93
    2017-05-01 00:32:55  30 -3519  1106  49.94
    2017-05-01 00:33:25  30 -3514  1105  49.10
    2017-05-01 00:33:55  30 -3517  1105  49.91
    2017-05-01 00:34:25  30 -3515  1104  49.94
    
    
    print(dict_of_dfs[2])
                          A    B     C      D
    2017-05-01 00:32:25  30  375  1643  49.93
    2017-05-01 00:32:55  30  372  1643  49.94
    2017-05-01 00:33:25  30  371  1643  49.10
    2017-05-01 00:33:55  30  377  1646  49.91
    2017-05-01 00:34:25  30  372  1645  49.94
    

    【讨论】:

    • 这很棒,它适用于所有列表列,而无需指定它们。
    • 这在控制台中对我来说没问题,但是当我将它包含在我的脚本中时,我收到了这条消息 NameError: ("free variable 'type' referenced before assignment in enclosing scope", 'occurred at index duration') 知道为什么吗?
    猜你喜欢
    • 2017-12-12
    • 2018-11-25
    • 1970-01-01
    • 2016-09-06
    • 2018-07-14
    • 2019-05-07
    • 1970-01-01
    • 2019-07-18
    相关资源
    最近更新 更多