【问题标题】:Find number of non-unique column values within a group where a second column value is all the same在第二列值都相同的组中查找非唯一列值的数量
【发布时间】:2021-05-24 14:24:13
【问题描述】:

我有两张桌子。

Table 1:
id_a, 
id_b, 
id_t
Table 2: 
id_t, 
name

如果表2名称以a开头,我需要找出任何东西 有匹配的id_ts 也有匹配的id_as。 如果表 2 名称以b 开头,我需要找出任何一行 有匹配的 id_ts 也有匹配的id_bs。

我需要知道这些匹配发生了多少次。

表 1

id_a id_b id_t
1 0 123
1 0 123
2 0 123
0 4 456
0 4 456
0 5 456
0 5 456
0 5 456
0 6 456
0 7 456

表 2

id_t name
123 aaq
456 bws

所以在这个例子中,我想看到这样的结果

id_t name num_non_unique
123 aaq 1
456 bws 2

我当前的代码是这样的:

SELECT
    t2.id_t, t2.name, count(t1.*) AS num_non_unique
FROM
    Table 2 AS t2
    JOIN Table 1 as t1 ON t2.id_t = t1.id_t
WHERE
        (t2.name like 'a%' and t1.id_a in (SELECT id_a FROM t1 GROUP BY id_a, id_t HAVING count(*) > 1))
    OR  (t2.name like 'b%' AND t1.id_b IN (SELECT id_b FROM t1 GROUP BY id_b, id_t HAVING count(*) > 1))
GROUP BY t1.name, t1.id_t

这目前没有给我想要的结果。 使用这段代码,我似乎得到了 id_b 的所有可用行的计数,以及 id_a 的 1 + non_uniques 的计数(因此,如果有一个非唯一的,则值为 2,否则该列的值为 1)。

感谢任何帮助!

【问题讨论】:

  • 你能解释一下你的输出吗?怎么会是1 for id_ts 123 amd 2 for id_ts 456
  • 因为 id_t 123 有一个值在 id_a 中重复,而对于 id_t 456 有两个值在 id_b 中重复。
  • 已经为您的要求添加了答案

标签: sql database postgresql


【解决方案1】:

架构和插入语句:

 create table table_1(id_a int, id_b int, id_t int);
 insert into table_1 values(1,  0,  123);
 insert into table_1 values(1,  0,  123);
 insert into table_1 values(2,  0,  123);
 insert into table_1 values(0,  4,  456);
 insert into table_1 values(0,  4,  456);
 insert into table_1 values(0,  5,  456);
 insert into table_1 values(0,  5,  456);
 insert into table_1 values(0,  5,  456);
 insert into table_1 values(0,  6,  456);
 insert into table_1 values(0,  7,  456);
 
 create table table_2 (id_t int,    name varchar(50));
 insert into table_2 values(123,    'aaq');
 insert into table_2 values(456,    'bws');

查询:

 with case1 as 
 (
   select id_t, id_a
   from table_1
   group by id_t, id_a
   having count(*)=1
 ),
 case2 as
 (
   select id_t,id_b
   from table_1
   group by id_t,id_b
   having count(*)=1
 )
 select id_t,name, 
 (case when t2.name like 'a%' then (select count(*)  from case1 where case1.id_t=t2.id_t) when t2.name like 'b%' then (select count(*) from  case2 where case2.id_t=t2.id_t) end)num_non_unique
 from table_2 as t2

输出:

id_t name num_non_unique
123 aaq 1
456 bws 2

dbhere

【讨论】:

    【解决方案2】:

    如果我理解正确,您需要计数 - 对于 table_2 中的每一行 - table_1 中的相应行,其中 id_b 只出现一次。

    一种方法是:

    select t2.*, coalesce(cnt, 0)
    from table_2 t2 left join
         (select id_a, count(*) as cnt
          from (select id_a, id_b
                from table_1
                group by id_a, id_b
                having count(*) = 1
               ) t1
          group by id_a
         ) t1
         on t1.id_a = t2.id_a;
    

    【讨论】:

      【解决方案3】:

      试试这个代码:

      select name,id_ts,sum(count_) "num_non_unique" from 
      (
      select t1.id_a,t1.id_b,t1.id_ts,t2.name, count(*) "count_"
      from tab1 t1 
      inner join tab2 t2 on t1.id_ts=t2.id_t
      group by 1,2,3,4
      having count(*)=1
      )tab
      group by 1,2
      

      说明: 正如您在问题中提到的那样,一列值将始终基于名称起始字符相同。因此无需检查查询中名称的开头。

      我们将检查唯一行的计数并使用条件having count(*)=1 对其进行过滤。

      获得唯一行后,只需将其分组即可获得按name和id_ts分组的行数

      DEMO

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 2023-01-24
        • 1970-01-01
        • 1970-01-01
        • 2018-08-10
        • 1970-01-01
        • 2013-03-03
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多