【问题标题】:Improving performance of flattening pandas DataFrames into a list of dicts提高将 pandas DataFrames 扁平化为 dicts 列表的性能
【发布时间】:2019-07-04 16:54:47
【问题描述】:

我有一个时间序列字典作为pandas.DataFrame 对象,每个对象都有任意数量的列。

我想将每个 DataFrame 转换为一个 dict 列表(例如,[{"col1": "row1", "col2": "row2", ..}, {"col1": "row2", ..}, ..],然后按每个 dict 的时间戳值对它们进行排序(每个 DataFrame 中必须有时间戳)。

这是一个性能改进问题。下面的代码有效,但我试图找到最快的方法来做到这一点。

我知道这个问题可以并行化,但不确定它是否是最佳路线。

import pandas as pd
import numpy as np


def gen_random_df(rows):
    df = pd.DataFrame({'x': np.random.normal(rows), 'y': np.random.normal(rows), 'z': np.random.normal(rows)},
                      index=pd.date_range('1900-01-01', '2049-12-31')[:rows])
    df.index.name = 'timestamp'
    return df


def to_list1(df, symbol):
    df = df.reset_index()
    return [dict(zip(df.columns, v), symbol=symbol) for v in df.values]


def method1(dict_of_dfs):
    data = []
    for symbol, df in dict_of_dfs.items():
        data.extend(to_list1(df, symbol))
    return sorted(data, key=lambda x: x['timestamp'])

第二种方法:


def method2(dict_of_dfs):
    dict_of_dfs = {symbol: df.assign(symbol=symbol) for symbol, df in dict_of_dfs.items()}
    data = pd.concat(dict_of_dfs.values(), axis=0).reset_index().to_dict('index').values()
    return list(data)

这是两种方法的性能。 Method1 是最快的,但可以改进吗?

symbols = 10
rows = 10_000
dict_of_dfs = {str(symbol): gen_random_df(rows) for symbol in range(symbols)}

%timeit result = method1(dict_of_dfs)
1.46 s ± 64.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
it
%timeit result = method2(dict_of_dfs)
1.87 s ± 102 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

这是预期的结果:

result[:3]
[{'timestamp': Timestamp('1900-01-01 00:00:00'),
  'x': 9998.31375178033,
  'y': 10000.298442533112,
  'z': 9999.538765089255,
  'symbol': '0'},
 {'timestamp': Timestamp('1900-01-02 00:00:00'),
  'x': 9998.31375178033,
  'y': 10000.298442533112,
  'z': 9999.538765089255,
  'symbol': '0'},
 {'timestamp': Timestamp('1900-01-03 00:00:00'),
  'x': 9998.31375178033,
  'y': 10000.298442533112,
  'z': 9999.538765089255,
  'symbol': '0'}]

【问题讨论】:

    标签: python pandas performance numpy numba


    【解决方案1】:

    基于this answer,我假设to_list1 的最快方法不是使用dict,而是使用chain 的dict 理解来迭代扩展值列表以及准备列名列表( cols) 提前。

    def to_list1(df, symbol):
        df = df.reset_index()
        cols = list(df.columns)
        cols.append('symbol')
    
        return [{kk:vv for kk,vv in zip(cols, chain(v, [symbol,]))} for v in df.values]
    

    就我而言(Python 3.7.2 64b Ubuntu 16.04)timeit 返回:

    to_list1: 2.211 s
    to_list2: 6.629 s
    

    【讨论】:

      猜你喜欢
      • 2014-02-23
      • 2015-03-11
      • 2021-05-24
      • 2021-03-06
      • 2019-04-23
      • 1970-01-01
      • 2016-01-14
      • 1970-01-01
      • 2018-05-14
      相关资源
      最近更新 更多