【发布时间】:2015-09-16 03:38:20
【问题描述】:
我需要 SAS 大师的建议 :)。
假设我有两个大数据集。第一个是一个巨大的数据集(大约 50-100Gb!),其中包含电话号码。第二个包含前缀(20-40 千次观察)。
我需要为每个电话号码的第一个表添加最合适的前缀。
例如,如果我有一个电话号码 +71230000 和前缀
+7
+71230
+7123
最合适的前缀是+71230。
我的想法。首先,对前缀表进行排序。然后在数据步骤中,处理电话号码表
data OutputTable;
set PhoneNumbersTable end=_last;
if _N_ = 1 then do;
dsid = open('PrefixTable');
end;
/* for each observation in PhoneNumbersTable:
1. Take the first digit of phone number (`+7`).
Look it up in PrefixTable. Store a number of observation of
this prefix (`n_obs`).
2. Take the first TWO digits of the phone number (`+71`).
Look it up in PrefixTable, starting with `n_obs + 1` observation.
Stop when we will find this prefix
(then store a number of observation of this prefix) or
when the first digit will change (then previous one was the
most appropriate prefix).
etc....
*/
if _last then do;
rc = close(dsid);
end;
run;
我希望我的想法足够清楚,但如果不是,我很抱歉)。
那你有什么建议? 感谢您的帮助。
附:当然,第一个表中的电话号码不是唯一的(可能重复),不幸的是,我的算法没有使用它。
【问题讨论】:
-
由于这个问题是针对 SAS 专家的,并且还涉及到大数据,所以我想知道你们是否可以帮我回答这个关于 SAS 数据存储选项和大数据的一般性问题:datascience.stackexchange.com/questions/12619/…。我正在尝试将 SAS 数据存储与 SQL Server 等常规 RDBMS 进行比较。对此的任何帮助都非常感谢。
标签: sas