【问题标题】:Restructure table by removing NULL values通过删除 NULL 值来重构表
【发布时间】:2018-04-30 09:17:51
【问题描述】:

我在 SQL 中有一个如下所示的表:

Customer   Product    1999     2000     2001    2002      2003
Smith      51         NULL     NULL       15      14      NULL
Jones      14           11        7     NULL    NULL      NULL
Jackson    13         NULL     NULL     NULL       3         9

每年一栏下的数字是金额,以美元为单位。每个客户有连续两年的金额,其余年份为零。我想重新构建此表,以便它只有两列 Amount-Year1 和 Amount-Year2,而不是年份列表。因此,它选择两个非零年份并将它们以正确的顺序放在这些列中。这将大大减小我的表的大小。

到目前为止,我已经能够对其进行重组,以便有一个金额列和一年列,但我随后会为每个客户获得多行,不幸的是我不能拥有(由于下游分析)。谁能想到获得两个 Amount-Year 列的方法?

我希望决赛桌看起来像这样:

Customer    Product    Amount_Y1    Amount_Y2
Smith       51         15           14
Jones       14         11            7
Jackson     13          3            9

我不介意丢失有关特定年份的信息,因为我可以从其他来源获得。实际表中有 1999 年至 2018 年间所有年份的数据,未来还会有更多年份。

谢谢

【问题讨论】:

  • 您使用的是哪个 RDBMS?
  • 对不起,应该说 - 我正在使用 SQL Server Management Studio
  • 如果我需要重新构建这个表,我会分成 2 个表 - tblCust (CustID, CustName) 和 tblCustAmount (RowID, FK_CustID, Amount)。这将有助于未来的数据参考和报告。

标签: sql sql-server database data-structures


【解决方案1】:

感谢UNPIVOT 无论如何都删除了NULLs,所以我们可以使用UNPIVOT/ROWNUMBER(),PIVOT 来做到这一点:

declare @t table (Customer varchar(15),Product int,[1999] int,
                  [2000] int,[2001] int,[2002] int,[2003] int)
insert into @T(CUstomer,Product,[1999],[2000],[2001],[2002],[2003]) values
('Smith'  ,51,NULL,NULL,  15,  14,NULL),
('Jones'  ,14,  11,   7,NULL,NULL,NULL),
('Jackson',13,NULL,NULL,NULL,   3,   9)

;With Numbered as (
    select
        Customer,Product,Value,
        ROW_NUMBER() OVER (PARTITION BY Customer,Product
                           ORDER BY Year) rn
    from
        @t t
            unpivot
        (Value for Year in ([1999],[2000],[2001],[2002],[2003])) u
)
select
    *
from
    Numbered n
        pivot
    (SUM(Value) for rn in ([1],[2])) w

结果:

Customer        Product     1           2
--------------- ----------- ----------- -----------
Jackson         13          3           9
Jones           14          11          7
Smith           51          15          14

【讨论】:

    【解决方案2】:

    使用COALESCE 将为您完成这项工作。查询是动态的,因此如果明天的列发生更改,即删除或添加,您无需更改任何内容。

    示例查询:(假设表为table1,列名与年份相同)。

        DECLARE @columnsdesc nvarchar(max), @columnsasc nvarchar(max)
    SET @columnsdesc = ''
    SELECT @columnsdesc = (select + '[' +  ltrim(c.Name)  +   ']' + ','
    FROM     sys.columns c 
             JOIN sys.objects o ON o.object_id = c.object_id 
    WHERE    o.type = 'U' and o.Name = 'table1' and c.Name not in ('Customer', 'Product')
    ORDER BY c.Name desc for xml path ( '' ))
    
    SET @columnsasc = ''
    SELECT @columnsasc = (select + '[' +  ltrim(c.Name)  +   ']' + ','
    FROM     sys.columns c 
             JOIN sys.objects o ON o.object_id = c.object_id 
    WHERE    o.type = 'U' and o.Name = 'table1' and c.Name not in ('Customer', 'Product')
    ORDER BY c.Name asc for xml path ( '' ))
       SELECT @columnsasc = LEFT( @columnsasc,LEN(@columnsasc)-1)
       SELECT @columnsdesc = LEFT( @columnsdesc,LEN(@columnsdesc)-1)
       DECLARE @sql nvarchar(max)
    
    SET @sql = 'SELECT Customer, Product, COALESCE('+ @columnsasc  +') as Amount_Y1,
    COALESCE(' + @columnsdesc +' ) as Amount_Y2
    FROM Table1'
    
    EXEC(@sql)
    

    如果您处理的是temporary table,那么代码会略有变化: 在这里测试:http://rextester.com/MRVR48808

    DECLARE @columnsdesc nvarchar(max), @columnsasc nvarchar(max)
    SET @columnsdesc = ''
    SELECT @columnsdesc = (select + '[' +  ltrim(c.Name)  +   ']' + ','
    FROM     tempdb.sys.columns c               --Changes here
             JOIN tempdb.sys.objects o ON o.object_id = c.object_id    --Changes here
    WHERE    o.type = 'U' and o.Name like '#table1%' and c.Name not in ('Customer', 'Product')   --Changes here
    ORDER BY c.Name desc for xml path ( '' ))
    
    SET @columnsasc = ''
    SELECT @columnsasc = (select + '[' +  ltrim(c.Name)  +   ']' + ','
    FROM     tempdb.sys.columns c                   --Changes here
             JOIN tempdb.sys.objects o ON o.object_id = c.object_id     --Changes here 
    WHERE    o.type = 'U' and o.Name like '#table1%' and c.Name not in ('Customer', 'Product')  --Changes here
    ORDER BY c.Name asc for xml path ( '' ))
       SELECT @columnsasc = LEFT( @columnsasc,LEN(@columnsasc)-1)
       SELECT @columnsdesc = LEFT( @columnsdesc,LEN(@columnsdesc)-1)
       DECLARE @sql nvarchar(max)
    
    SET @sql = 'SELECT Customer, Product, COALESCE('+ @columnsasc  +') as Amount_Y1,
    COALESCE(' + @columnsdesc +' ) as Amount_Y2
    FROM #Table1'   --Changes here
    
    EXEC(@sql)
    

    【讨论】:

    • 我真的很喜欢动态查询来获取列名的想法,但不幸的是我无法让代码工作。我认为您不能在派生表上使用 ORDER BY?
    • 我听到了。我已经进行了必要的更改并对其进行了测试,它工作正常。或者,您可以在这里进行测试:rextester.com/YGNE23015
    • 忘了提及,我还编辑了答案以包含更改。
    • 至少追溯到 2000 版本的 SQL Server,this warning:“不要在 SELECT 语句中使用变量来连接值(即计算聚合值)。意外查询结果可能会发生。这是因为 SELECT 列表中的所有表达式(包括赋值)不能保证对每个输出行都执行一次。我长期以来一直认为它在错误的页面上,但它确实表明不应该依赖这种技术。
    • 谢谢@Damien_The_Unbeliever。老实说,我在使用早期的方法时从来没有遇到过问题,但很高兴你提出来了。我已经进行了进一步所需的更改以避免该问题并对其进行了测试。参考:Technet 文章以安全地连接值并避免上述问题:social.technet.microsoft.com/wiki/contents/articles/… rw2:请试一试,让我知道它是如何进行的。另外,我在这里更新了它:rextester.com/edit/YGNE23015
    【解决方案3】:

    尝试按如下方式使用 COALESCE:一个字段从头到尾,第二个字段则相反。

    SELECT Customer,Product, COALESCE([1999],[2000],[2001],[2002],[2003]) as Y1, 
           COALESCE([2003],[2002],[2001],[2000],[1999])  as Y2
    FROM  #TEMPDATA
    

    【讨论】:

    • Aswani - 你能把这个动态化吗,因为可能有 3 个非空值。
    • @Pawan Kumar 那样的话,三选二的标准是什么?
    • 试试这个例子 - ('Jackson',NULL,NULL,NULL,NULL,5)
    【解决方案4】:

    我会使用cross apply:

    select t.customer, t.product, v.Amount_Y1, v.Amount_Y2
    from t cross apply
         (select max(case when which = 1 then val end) as Amount_Y1,
                 max(case when which = 2 then val end) as Amount_Y2
          from (select val, yr, row_number() over (order by yr) as which
                from (values (t.[1999], 1999), (t.[2000], 2000), (t.[2001], 2001),
                             (t.[2002], 2002), (t.[2003], 2003)
                     ) v(val, yr)
                where val is not null
               ) v
    

    【讨论】:

      猜你喜欢
      • 2016-04-29
      • 1970-01-01
      • 2021-08-17
      • 2018-07-03
      • 1970-01-01
      • 2013-12-22
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多