【问题标题】:pandas_profiling taking way too long to runpandas_profiling 运行时间过长
【发布时间】:2020-04-16 10:16:26
【问题描述】:

如果有人尝试过pandas-profiling package,请帮助我提供任何关于使其运行更快的见解。包中的输出报告非常整洁和详细,但是即使使用中等大小的数据集,创建报告也需要很长时间。来自 Kaggle 推土机数据集的大约 10 列和 400K 行耗时 21 分钟(非 GPU)。想知道它是否值得进一步研究。

df.shape
(401125, 9)


start = datetime.datetime.now()
profile = df.profile_report(title="Exploring Dataset")
profile.to_file(output_file=Path("./data_report.html"))

end = datetime.datetime.now()
print(end-start)

0:21:23.976324

【问题讨论】:

    标签: pandas pandas-profiling


    【解决方案1】:

    根据您的兴趣,您可以禁用 pandas-profiling 的其他耗时最多的功能,因为它是模块化的。这是目前您加速以及对数据集进行采样的首选解决方案。

    这里有几个相关的问题:

    从长远来看,我们计划允许更好的并行化和更合理的默认值: https://github.com/pandas-profiling/pandas-profiling/issues/279

    编辑:

    从 v2.4 开始,存在最小模式,可将包配置为自动使用低计算设置:https://github.com/pandas-profiling/pandas-profiling#large-datasets

    【讨论】:

      猜你喜欢
      • 2018-04-07
      • 2016-06-25
      • 1970-01-01
      • 1970-01-01
      • 2013-05-12
      • 2012-07-13
      • 2018-07-02
      • 2011-11-11
      • 1970-01-01
      相关资源
      最近更新 更多