【发布时间】:2021-03-25 06:41:12
【问题描述】:
执行操作时:Dask.dataframe.to_parquet(data),如果通过Dask 以给定数量的分区读取data,并且您在删除一些列后尝试以镶木地板格式保存它,它会失败,例如以下错误:
FileNotFoundError: [Errno 2] No such file or directory: part.0.parquet'
有人遇到过同样的问题吗?
这是一个最小的示例 - 请注意方式 1 可以按预期工作,而方式 2 则不能:
import numpy as np
import pandas as pd
import dask.dataframe as dd
# -------------
# way 1 - works
# -------------
print('way 1 - start')
A = np.random.rand(200,300)
cols = np.arange(0, A.shape[1])
cols = [str(col) for col in cols]
df = pd.DataFrame(A, columns=cols)
ddf = dd.from_pandas(df, npartitions=11)
# compute and resave
ddf.drop(cols[0:11], axis=1)
dd.to_parquet(
ddf, 'error.parquet', engine='auto', compression='default',
write_index=True, overwrite=True, append=False)
print('way 1 - end')
# ----------------------
# way 2 - does NOT work
# ----------------------
print('way 2 - start')
ddf = dd.read_parquet('error.parquet')
# compute and resave
ddf.drop(cols[0:11], axis=1)
dd.to_parquet(
ddf, 'error.parquet', engine='auto', compression='default',
write_index=True, overwrite=True, append=False)
print('way 2 - end')
【问题讨论】:
标签: python python-3.x dask parquet dask-dataframe