【问题标题】:How to extract values from a string in BigQuery and assign a position number?如何从 BigQuery 中的字符串中提取值并分配位置编号?
【发布时间】:2017-04-21 05:55:56
【问题描述】:

我的数据目前如下所示:

我想要的输出是这样的:

期望的成就是:

  • 在 OrderDescription 中提取 .csv 字符串中的值。
  • 将它们呈现在一个规范化的表中,并分配 OrderDescriptionPosition,即原始数据中 .csv 字符串中呈现的顺序。

我通过结合使用 split 函数和 row_number 来做到这一点。这似乎在逐个客户的基础上工作,但在运行多个客户时似乎返回不可靠的“洗牌”row_numbers,即不正确的订单。

select
   CustomerID,
   OrderID,
   OrderDescriptionItem,            
   row_number() over(partition by CustomerID, OrderID) as OrderDescriptionPosition
from
(
    select
       CustomerID,
       OrderID,          
       split(OrderDescription, ',') as OrderDescriptionItem
    from
       InitialTable
) as e
,unnest(OrderDescriptionItem)   as OrderDescriptionItem

有没有人有更强大的解决方案?欢迎使用 UDF 和 javascript 的任何建议。

【问题讨论】:

    标签: google-bigquery


    【解决方案1】:

    您可以将WITH OFFSET 与UNNEST 结合使用来获取职位。这是一个例子:

    #standardSQL
    WITH Input AS (
      SELECT 1 AS CustomerID, 1001 AS OrderID, '12,14,16,22,28' AS OrderDescription UNION ALL
      SELECT 2 AS CustomerID, 1002 AS OrderID, '1,5' AS OrderDescription UNION ALL
      SELECT 3 AS CustomerID, 1003 AS OrderID, '44,55,66' AS OrderDescription
    )
    SELECT
      CustomerID,
      OrderID,
      OrderDescription,
      off + 1 AS OrderDescriptionPosition
    FROM Input
    CROSS JOIN UNNEST(SPLIT(OrderDescription)) AS OrderDescription
      WITH OFFSET off;
    +------------+---------+------------------+--------------------------+
    | CustomerID | OrderID | OrderDescription | OrderDescriptionPosition |
    +------------+---------+------------------+--------------------------+
    | 1          | 1001    | 12               | 1                        |
    | 1          | 1001    | 14               | 2                        |
    | 1          | 1001    | 16               | 3                        |
    | 1          | 1001    | 22               | 4                        |
    | 1          | 1001    | 28               | 5                        |
    | 2          | 1002    | 1                | 1                        |
    | 2          | 1002    | 5                | 2                        |
    | 3          | 1003    | 44               | 1                        |
    | 3          | 1003    | 55               | 2                        |
    | 3          | 1003    | 66               | 3                        |
    +------------+---------+------------------+--------------------------+
    

    【讨论】:

      【解决方案2】:

      如果您的示例代表您的真实用例(在某种意义上 OrderDescription 是值的有序列表) - 您几乎可以按原样使用您的查询版本 - 只需在 OVER() 中添加 ORDER BY如下

      #standardSQL
      WITH InitialTable AS (
        SELECT 1 AS CustomerID, 1001 AS OrderID, '12,14,16,22,28' AS OrderDescription UNION ALL
        SELECT 2, 1002, '1,5' UNION ALL
        SELECT 3, 1003, '44,55,66'
      )
      SELECT
        CustomerID,
        OrderID,
        OrderDescription,
        ROW_NUMBER() OVER(PARTITION BY CustomerID, OrderID ORDER BY OrderDescription) AS OrderDescriptionPosition
      FROM InitialTable, UNNEST(SPLIT(OrderDescription)) AS OrderDescription
      -- ORDER BY CustomerID, OrderID, OrderDescriptionPosition
      

      【讨论】:

      • 感谢您的反馈米哈伊尔。也许我应该在我的示例中澄清“OrderDescription”的顺序并不总是升序,并且 .csv 字符串中的位置非常重要。 OrderDescription = (6, 1, 400, 43) 应该返回两列:OrderDescription [6, 1, 400, 43] 和 OrderDescriptionPosition [1, 2, 3, 4]
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2019-05-31
      • 1970-01-01
      • 2017-07-22
      • 1970-01-01
      • 2018-12-10
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多