【问题标题】:Drop all observations by ID where conditions are not met在不满足条件的情况下按 ID 删除所有观察值
【发布时间】:2015-05-16 16:11:54
【问题描述】:

我有一个包含约 400 万笔交易记录的数据集,按 Customer_No 分组(每个 Customer_No 包含 1 个或多个交易,用顺序计数器表示)。每笔交易都有一个类型代码,我只对使用特定交易类型组合的客户感兴趣。加入表本身或在 Proc Sql 中使用 EXISTS 都不允许我有效地评估事务类型标准。我怀疑使用保留和执行循环的数据步骤会更快地处理数据集

数据集:

Customer_No Tran_Seq    Tran_Type
    0001        1           05
    0001        2           12
    0002        1           07
    0002        2           86
    0002        3           04
    0003        1           07
    0003        2           84
    0003        3           84
    0003        4           84

我尝试应用的标准:

  1. 所有 Customer_No 的 Tran_Type 只能在 ('04','05','07','84','86') 中, 如果使用了任何其他 Tran_Type,则删除该 Customer_No 的所有事务

  2. Customer_No 的 Tran_Type 必须包括('84' 或 '86')AND '04',如果不满足此条件,则删除 Customer_No 的所有事务

我想要的输出:

Customer_No Tran_Seq    Tran_Type
0002        1           07
0002        2           86
0002        3           04  

【问题讨论】:

  • 是您需要考虑的唯一代码还是还有更多?即这是对实际问题的简化吗?
  • 我会使用retain 语句设置指标以跨ID 保存它们,评估last.ID 的状态,然后输出ID 列表。但这不是代码编写服务,因此您需要先尝试一下:)。
  • 这是一种简化。实际代码为 3 个字符,共有 122 个字符。谢谢,我会试试你的方法
  • 我认为 proc sql 可以很好地解决这个问题。您的查询结构是什么?

标签: sas


【解决方案1】:

如果对数据进行排序,DoW 循环解决方案应该是最有效的。如果它没有排序,它要么是最有效的,要么是规模相似但效率稍低的,具体取决于数据集的情况。

我将 Dom 的解决方案与 3e7 ID 数据集进行了比较,得到的 DoW 总长度相似(略短),未排序数据集的 CPU 更少,排序后的速度大约快 50%。保证在大约写出数据集所需的时间长度内运行(可能会多一点,但应该不会太多),如果需要,再加上排序时间。

data want;
  do _n_=1 by 1 until (last.customer_no);
      set have;
      by customer_no;  
      if tran_type in ('84','86') 
        then has_8486 = 1;
      else if tran_type in ('04') 
        then has_04 = 1;
      else if not (tran_type in ('04','05','07','84','86')) 
        then has_other = 1;
  end;
  do _n_= 1 by 1 until (last.customer_no);
    set have;
    by customer_no;
    if has_8486 and has_04 and not has_other then output;
  end;
run;

【讨论】:

    【解决方案2】:

    我认为没有那么复杂。加入子查询group by Customer_No,并将您的条件放在having 子句中。 min 函数中的条件对于所有行都必须为真,而max 函数中的条件对于任何一行都必须为真:

    proc sql;
    create table want as
    select
      h.*
    from
      have h
      inner join (
        select
          Customer_No
        from
          have
        group by
          Customer_No
        having
          min(Tran_Type in('04','05','07','84','86')) and
          max(Tran_Type in('84','86')) and
          max(Tran_Type eq '04')) h2
      on h.Customer_No = h2.Customer_No
    ;
    quit;
    

    【讨论】:

    • 我喜欢这种方法,但认为您可以完全不使用连接和子查询。只需select * from have group by Customer_No having ...。
    • 你是对的,它会产生相同的结果。我故意避免这种方法有两个原因:(1)它不像明确地进行连接那么清楚; SQL 的其他迭代不允许这种类型的快捷方式(我相信这是有充分理由的)。 (2) 根据我的经验,当您依靠“重新合并汇总统计”方法时,在 SAS 中运行似乎需要更长的时间。我不知道为什么它的效率较低,但这是我的经验。
    【解决方案3】:

    我会使用 INTERSECT 运算符提供比 @naed555 稍微简单的 SQL 解决方案。

    proc sql noprint;
    
    create table to_keep as
    (
        select distinct customer_no
        from have 
        where tran_type in ('84','86')
    
        INTERSECT
    
        select distinct customer_no
        from have 
        where tran_type in ('04')
    )
    
    EXCEPT
    
        select distinct customer_no
        from have
        where tran_type not in ('04','05','07','84','86')
    ;
    
    create table want as
    select a.*
    from have as a
    inner join 
         to_keep as b
    on a.customer_no = b.customer_no;
    
    quit;
    

    【讨论】:

      【解决方案4】:

      我一定犯了一个加入错误。重写时,Proc Sql 在不到 30 秒的时间内完成(在原始的 490 万条记录数据集上)。虽然它不是特别优雅的代码,所以我仍然希望有任何改进或替代方法。

      data Have;
      input Customer_No $ Tran_Seq $ Tran_Type:$2.;
      cards;
          0001        1           05
          0001        2           12
          0002        1           07
          0002        2           86
          0002        3           04
          0003        1           07
          0003        2           84
          0003        3           84
          0003        4           84
      ;
      run;
      
      Proc sql;
      Create table Want as
      select t1.* from Have t1
      LEFT JOIN (select DISTINCT Customer_No from Have
                          where Tran_Type not in ('04','05','07','84','86')
                                        ) t2
      ON(t1.Customer_No=t2.Customer_No)
      INNER JOIN (select DISTINCT Customer_No from Have
                          where Tran_Type in ('84','86')
                                      ) t3
      ON(t1.Customer_No=t3.Customer_No)
      INNER JOIN (select DISTINCT Customer_No from Have
                          where Tran_Type in ('04')
                                      ) t4
      ON(t1.Customer_No=t4.Customer_No)
      Where t2.Customer_No is null
      ;Quit;
      

      【讨论】:

      • 我的回答基本上是这样做的,但看起来更干净一些。我不确定哪个会跑得最快。如果可以,请尝试在 TRAN_TYPE 上创建索引,因为它可能有助于子查询。
      • 我认为您的代码中存在逻辑问题。第一部分的 LEFT JOIN 不会过滤掉 Customer 001。你得到的表是正确的,但那是因为它后来被过滤掉了。当您的客户通过 (84,86) 和 (04) 测试但第一次未通过时,这将导致较大的表出现问题。
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2016-01-29
      • 2018-10-24
      • 1970-01-01
      • 1970-01-01
      • 2020-02-09
      • 1970-01-01
      • 2016-12-04
      相关资源
      最近更新 更多