【问题标题】:Select three columns (A,B,C) and return rows where (B,C) are distinct and A is at max value because there are multiple occurrences of (B,C) pairs选择三列 (A,B,C) 并返回 (B,C) 不同且 A 处于最大值的行,因为 (B,C) 对出现多次
【发布时间】:2017-06-30 07:49:18
【问题描述】:

在 Hive 中工作,我有一个包含三列 A(INT)、B(STRING) 和 C(STRING) 的表。 B 和 C 列具有重复的配对值(即第 1 行和第 10 行可能在 B 和 C 列中具有相同的字符串)。我正在尝试返回不同 B、C 对的完整行(即 A、B、C),其中 A 列在所有出现的不同 B、C 对中处于最大值。感谢所有帮助。

输入表示例

Col1 Col2 Col3
----------------
1111, str1, str2 
2222, str1, str2
3333, str3, str4
4444, str5, str6
5555, str3, str4
6666, str5, str6

查询输出示例

Col1 Col2 Col3
----------------
2222, str1, str2
5555, str3, str4
6666, str5, str6

【问题讨论】:

    标签: sql hive max distinct


    【解决方案1】:

    如果您只想要结果中的三列:

    select max(a) a, b, c
    from your_table
    group by b, c;
    

    如果你有更多的列要选择,你可以使用窗口函数row_number

    select *
    from (
        select t.*,
            row_number() over (
                partition by b, c
                order by a desc nulls last
                ) rn
        from your_table t
        ) t
    where rn = 1;
    

    【讨论】:

    • 我使用了第一个更简单的建议。将表从 54M 行重复缩小到 1.6M 唯一行。第二个建议返回 NULLS。我检查了输入表,任何列中都没有 NULLS。谢谢@GurV
    猜你喜欢
    • 1970-01-01
    • 2011-08-01
    • 2021-10-04
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2015-06-12
    • 2020-07-15
    • 2011-05-30
    相关资源
    最近更新 更多