【问题标题】:How to concatenate dask Dataframes with datetime index faster?如何更快地将 dask Dataframes 与日期时间索引连接起来?
【发布时间】:2019-01-24 09:53:01
【问题描述】:

在垂直连接两个时间戳索引的 dask Dataframe 时,我遇到了与 this 类似的问题。

我有两个 dask 数据框 df1,df2:

df1.index:

Dask Index Structure:

npartitions=1

2018-03-03 13:04:44.497929    datetime64[ns]

2018-03-03 13:23:04.759840               ...

Name: time, dtype: datetime64[ns]

Dask Name: getitem, 8 tasks

df2.index:

Dask Index Structure:

npartitions=1

2018-03-03 07:09:04.184453    datetime64[ns]

2018-03-03 07:32:46.815356               ...

Name: time, dtype: datetime64[ns]

Dask Name: getitem, 8 tasks

它们具有完全相同的列名和类型。现在我想使用 dask.dataframe.concat 连接它们:

#df1 & df2 are dask dataframes

print(df1.divisions)

print(df2.divisions)

dfs=dd.concat([df1,df2],axis=0,interleave_partitions=False)

输出:

(时间戳('2018-03-03 13:04:44.497929'), 时间戳('2018-03-03 13:23:04.759840')) (时间戳('2018-03-03 07:09:04.184453'),时间戳('2018-03-03 07:32:46.815356')) ValueError:所有输入都有无法按顺序连接的已知除法。指定 interleave_partitions=True 忽略顺序


两个 ddf 不能连接,除非指定 interleave_partitions=True。但是两个数据帧的索引之间没有交错。是不是dask支持datetimeindex的限制造成的?或者我需要指定其他参数或将索引转换为int或double?

【问题讨论】:

    标签: python pandas dask


    【解决方案1】:

    但是两个数据帧的索引之间没有交错

    Dask 在这里似乎不同意你的看法。似乎认为您的两个数据框的索引范围确实有些重叠。没关系,你可以按要求添加关键字,应该没问题。

    dfs=dd.concat([df1,df2],axis=0,interleave_partitions=True)
    

    如果您认为自己在这里遇到了错误,那么我鼓励您将其缩减为最小示例并发布错误报告。

    【讨论】:

    • 非常感谢。我从属性“ddf.divisons”中得到“两个数据帧之间没有交错”的结论。所以我认为当分区显示没有交错时,dask 数据帧不会重叠(每个数据帧的 npatitions=1,如上面的输出)。我是否从 ddf.divisions 输出中得出了错误的结论?
    • 您说得对,.divisions 属性完全确定是否引发该警告。
    猜你喜欢
    • 2019-09-15
    • 1970-01-01
    • 1970-01-01
    • 2013-11-15
    • 2015-09-28
    • 2019-10-09
    • 2020-06-11
    • 2016-05-18
    • 1970-01-01
    相关资源
    最近更新 更多