【问题标题】:Identifying first occurrence after trigger event识别触发事件后的第一次发生
【发布时间】:2018-07-09 15:13:29
【问题描述】:

我有一个看起来有点像这样的大面板数据集:

data have;
   input id t a b ;
datalines;
1 1 0 0
1 2 0 0
1 3 1 0
1 4 0 0
1 5 0 1
1 6 1 0
1 7 0 0
1 8 0 0
1 9 0 0
1 10 0 1
2 1 0 0
2 2 1 0
2 3 0 0
2 4 0 0
2 5 0 1
2 6 0 1
2 7 0 1
2 8 0 1
2 9 1 0
2 10 0 1
3 1 0 0
3 2 0 0
3 3 0 0
3 4 0 0
3 5 0 0
3 6 0 0
3 7 1 0
3 8 0 0
3 9 0 0
3 10 0 0
;
run;

对于每个 ID,我想记录所有“触发”事件,即当 a=1 时,然后我需要多长时间才能 下一次 出现 b=1。最终输出应该给我以下内容:

data want;
  input id a_no a_t b_t diff ;
datalines;
1 1 3 5 2
1 2 6 10 4
2 1 2 5 3
2 2 9 10 1
3 1 7 . .
;
run;

获取所有 a=1 和 b=1 事件当然没有问题,但由于它是一个非常大的数据集,每个 ID 都有很多这两个事件,我正在寻找一个优雅而直接的解决方案。有什么想法吗?

【问题讨论】:

  • 是否存在 a=1 和 b=1 在同一 t 的情况?如果是这样,应该怎么办?
  • a=1 和 b=1 不能同时发生

标签: sas panel-data


【解决方案1】:

这是一个相当简单的 SQL 方法,它或多或少地提供了所需的输出:

proc sql;
create table want
  as select 
    t1.id, 
    t1.t as a_t, 
    t2.t as b_t, 
    t2.t - t1.t as diff
    from 
      have(where = (a=1)) t1 
      left join 
      have(where = (b=1)) t2
    on 
      t1.id = t2.id 
      and t2.t > t1.t
    group by t1.id, t1.t
    having diff = min(diff)
    ;
quit;

唯一缺少的部分是a_no。这种行增量 ID 需要在 SQL 中始终如一地生成大量工作,但对于额外的数据步骤来说是微不足道的:

data want;
 set want;
 by id;
 if first.id then a_no = 0;
 a_no + 1;
run;

【讨论】:

  • 效果很好!我总是在 proc sql 上苦苦挣扎,所以这个解决方案之前并没有让我印象深刻。谢谢!
【解决方案2】:

优雅的 DATA 步方法可以使用嵌套的 DOW 循环。当您了解 DOW 循环时,这很简单。

data want(keep=id--diff);
  length id a_no a_t b_t diff 8;
  do until (last.id);                           * process each group;
    do a_no = 1 by 1 until(last.id);            * counter for each output;
      do until ( output_condition or end);      * process each triggering state change;

        SET have end=end;          * read data;
        by id;                     * setup first. last. variables for group;

        if a=1 then a_t = t;       * detect and record start of trigger state;

        output_condition = (b=1 and t > a_t > 0);  * evaluate for proper end of trigger state;
      end;

      if output_condition then do; 
        b_t = t;                     * compute remaining info at output point;
        diff = b_t - a_t;

        OUTPUT;

        a_t = .;       * reset trigger state tracking variables;
        b_t = .;
      end;
      else 
        OUTPUT;        * end of data reached without triggered output;
    end;
  end;
run;

注意:SQL 方式(未显示)可以在组内使用自联接。

【讨论】:

  • 我认为自己是一个 DOW 循环爱好者,但我发现这很难理解。此外,您可以将 (where = (a=1 or b=1)) 添加到您的 set 语句中。
  • 深度 DOW 循环处理的一个好处是发生源数据集的单次传递。找出从 SQL(需要连接)到 DOW 的性能考虑提示的数据的“性质”会很有趣。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2014-12-09
  • 1970-01-01
相关资源
最近更新 更多