【问题标题】:Pandas merge_asof on multiple columns多列上的 Pandas merge_asof
【发布时间】:2018-11-05 15:24:24
【问题描述】:

我有两个数据框:

DF1:

StartDate      Location

2013-01-01     20000002
2013-03-01     20000002
2013-08-01     20000002
2013-01-01     20000003
2013-03-01     20000003
2013-05-01     20000003
2013-01-01     20000043

DF2:

EmpStartDate   Location

2012-12-17     20000002.0 
2013-02-25     20000002.0 
2013-06-26     20000002.0 
2012-09-24     20000003.0 
2013-01-07     20000003.0 
2013-07-01     20000043.0

我想要 DF2 的计数,其中 DF1.Location = DF2.Location 和 DF2.EmpStartDate

输出:

StartDate      Location   Count

2013-01-01     20000002   1
2013-03-01     20000002   2
2013-08-01     20000002   3
2013-01-01     20000003   1
2013-03-01     20000003   2
2013-05-01     20000003   2
2013-01-01     20000043   0

我在 DF2.EmpStartDate 和 DF1.StartDate 上使用 merge_asof,然后在 Location 和 StartDate 上进行分组以实现此目的。 但是我得到的结果不正确,因为我只在日期列上合并。我需要合并 Location 和 Date 列上的数据框。看起来 merge_asof 不支持合并多个列。如何合并不同位置组的日期列?

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    merge_asof 保持left DataFrame 的大小,所以它不能将left 中的同一行匹配到right 中的多行。

    一种简单但内存效率低的计算方法是在Location 上执行一个大的merge,然后计算有多少行有df.EmpStartDate < df.StartDate

    df = df1.merge(df2)
    (df.assign(Count = df.EmpStartDate < df.StartDate)
       .groupby(['StartDate', 'Location'])
       .Count.sum()
       .astype('int')
       .reset_index())
    

    输出:

       StartDate  Location  Count
    0 2013-01-01  20000002      1
    1 2013-01-01  20000003      1
    2 2013-01-01  20000043      0
    3 2013-03-01  20000002      2
    4 2013-03-01  20000003      2
    5 2013-05-01  20000003      2
    6 2013-08-01  20000002      3
    

    【讨论】:

    • "merge_asof 只能产生 1:1 合并,所以我认为这不是您想要的。" ——是什么让你这么说?在很多情况下,操作可以对“左侧”数据框中的多行使用相同的数据?
    【解决方案2】:

    让我们使用这个:

    df1.merge(df2, on='Location')\
       .query('EmpStartDate <= StartDate')\
       .groupby(['StartDate','Location'])['EmpStartDate']\
       .count()\
       .reindex(df1, fill_value=0)\
       .rename('Count')\
       .reset_index()
    

    输出:

       StartDate  Location  Count
    0 2013-01-01  20000002      1
    1 2013-03-01  20000002      2
    2 2013-08-01  20000002      3
    3 2013-01-01  20000003      1
    4 2013-03-01  20000003      2
    5 2013-05-01  20000003      2
    6 2013-01-01  20000043      0
    

    【讨论】:

    • 我通过重新索引将计数设为 0。
    • 是的,那些缺少日期和位置,用 0 填充值。如果您不希望这样,则可以从重新索引中删除 'fill_value' 参数。
    • 我的意思是整个结果集都是 0!如果我删除重新索引,我会得到正确的结果。但是,如果有任何缺失值,我将不得不处理它..
    • 我不明白。如果您删除重新索引,您会得到正确的结果吗?要获取丢失的位置/日期,您需要生成所有可能的位置和日期的列表,然后使用重新索引。
    • 是的,如果我删除重新索引,我会得到正确的结果。如果我包括重新索引,则所有行的计数都为 0。目前,数据中没有遗漏地点或日期。所以,我猜它正在重新索引所有行。但是,每个月的数据可能会有所不同。所以,我将不得不处理它并编写一个通用代码。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-03-14
    • 2021-08-28
    • 1970-01-01
    • 2021-06-24
    • 2018-08-08
    • 2022-11-01
    相关资源
    最近更新 更多