【问题标题】:Find duplicates within a column and pull into new df在列中查找重复项并拉入新的 df
【发布时间】:2023-03-31 18:16:01
【问题描述】:

我有两个数据框,其中一列中有重复项。我需要将重复项提取到单独的数据框中。

df1=
    col1     col2   col3   col4
    BLUE      .5     yes    5
    GREEN     .2     no     2
    PINK      .3     yes    3

df2=
    col1     col2   col3   col4
    RED       .9     yes    9
    GREEN     .2     yes    2
    BLUE      .7     yes    7

例如,虽然GREEN行的col2不一样,但我只需要从df2中提取出来,拉到一个新的df中即可。

但还要处理数千行以希望找到一种批量执行此操作的方法。

【问题讨论】:

    标签: python python-3.x duplicates


    【解决方案1】:

    IIUC 你只需要在col1 上使用.merge:

    new_df = df2.merge(df1['col1'], on='col1')
    

    输出:

        col1  col2 col3  col4
    0  GREEN   0.2  yes     2
    1   BLUE   0.7  yes     7
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2020-09-10
      • 2021-07-03
      • 2013-02-21
      • 2013-07-06
      相关资源
      最近更新 更多