【发布时间】:2020-04-16 10:16:26
【问题描述】:
如果有人尝试过pandas-profiling package,请帮助我提供任何关于使其运行更快的见解。包中的输出报告非常整洁和详细,但是即使使用中等大小的数据集,创建报告也需要很长时间。来自 Kaggle 推土机数据集的大约 10 列和 400K 行耗时 21 分钟(非 GPU)。想知道它是否值得进一步研究。
df.shape
(401125, 9)
start = datetime.datetime.now()
profile = df.profile_report(title="Exploring Dataset")
profile.to_file(output_file=Path("./data_report.html"))
end = datetime.datetime.now()
print(end-start)
0:21:23.976324
【问题讨论】: