【问题标题】:SQL: select unique rowsSQL:选择唯一行
【发布时间】:2021-07-03 19:06:18
【问题描述】:

这是一个表的“玩具”示例,它有很多列和数千行。

我想要过滤掉任何包含相同 AcctNo、CustomerName 和 CustomerContact 的行,但保留 ONE 重复项的 ID(这样我以后可以访问记录) .

  • 例子:

    ID  AcctNo  CustomerName  CustomerContact
    1   1111    Acme Foods    John Smith
    2   1111    Acme Foods    John Smith
    3   1111    Acme Foods    Judy Lawson
    4   2222    YoyoDyne Inc  Thomas Pynchon
    5   2222    YoyoDyne Inc  Thomas Pynchon
    <= I want to save IDs 2, 3, and 5
    
  • 小提琴: https://www.db-fiddle.com/f/bEECHi6XnvKAeXC4Xthrrr/1

问:我需要什么 SQL 来完成这个?

【问题讨论】:

  • 你试过什么?你在哪里卡住了?请向我们展示您的尝试。
  • ID 3 是怎么重复的?
  • 你可以考虑使用row_number()函数。
  • 请分享你已经尝试过的sql。
  • 您需要每个组的最大 ID...

标签: sql sql-server group-by distinct


【解决方案1】:
select MAX(ID) as KeepID,AcctNo,CustomerName,CustomerContact 
from test
GROUP BY AcctNo,CustomerName,CustomerContact

【讨论】:

    【解决方案2】:

    所以基本上你想要的是,按 AcctNo、CustomerName 和 CustomerContact 对表进行分区。问题中不清楚您希望如何选择需要保留的 ID,但为此您需要修改以下查询。但这应该给你一个起点。

    SELECT * 
    FROM   test 
           JOIN (SELECT id, 
                        Row_number() 
                          OVER ( 
                            partition BY acctno, customername, customercontact) rn 
                 FROM   test) A 
             ON test.id = A.id 
    WHERE  A.rn = 1 
    

    这应该返回如下内容:

    ID AcctNo CustomerName CustomerContact id rn
    1 11111 Acme Foods John Smith 1 1
    3 11111 Acme Foods Judy Lawson 3 1
    4 22222 Yoyodyne Inc. Thomas Pynchon 4 1

    这基本上是首先根据分区标准计算行数,然后每个分区只选择一行。

    【讨论】:

    • 请不要将图像用于数据...使用格式化/表格文本。
    • @Pradatta - 1)在回答您上面的评论时,我尝试了很多东西......但我不想用一堆失败的尝试“弄乱”我的问题。 2) 关于您的回复:我考虑过 分区和row_number(),但我更喜欢更简单的解决方案。喜欢Robert Sheahan's。 3) 你知道哪种方法更“有效”(对于大型数据集)?
    • @FoggyDay 我认为带有 Group By 的 MAX 效率更高。这是一个很好的解释:stackoverflow.com/questions/11233125/…
    • 优秀的引用 - 谢谢。
    猜你喜欢
    • 1970-01-01
    • 2013-02-14
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2022-12-09
    • 1970-01-01
    相关资源
    最近更新 更多