更新:不使用保留/恢复(感谢@Pearly Spencer 强调这一点)进一步提高了方法的速度。要查看保留/恢复的旧代码,请参阅the older versions of this answer。
我想我找到了一种更快的方法来解决问题(至少从使用timer on、timer off 的结果来看)。
因此,回顾一下,当前缓慢的方法是合并数据库,然后使用
删除所有未使用的标签
labelbook, problems
label drop `r(notused)'
另一种更快的方法是加载较小的数据集,仅使用所需的变量。这将仅包含所选变量的标签。然后,将这个较小的数据库与原始数据库合并。重要的是,合并方向是相反的!通过这种方式,我们消除了保留/恢复的需要,正如@Pearly Spencer 所建议的那样,这可能会减慢速度,尤其是在较大的数据集中。
就我原来的例子而言,代码是:
*** Open and work with dataset A ***
use A.dta // load original dataset
... // do stuff with it (added just for generality)
save A_final.dta // name of final dataset
*** Load dataset B with subset of needed variables only ***
use id var1 var2 using B.dta, clear // this loads id (needed for merging), var1 and var2 and their labels only!
*** Merge modified A dataset into smaller B dataset ***
merge 1:1 id using A_final.dta, keep(using match) // we do not specify any variables to load, as all those in A_final.dta needed) IMPORTANT: if we want to keep all observations of the original dataset (A, which is the one being merged into B), we need to use "using" rather than "master" in the "keep()" option.
save A_final.dta, replace // Create final version of A. Done!
就是这样!我不确定这是否是最佳解决方案,但在我的情况下,我正在合并具有数百个变量的许多数据集,它的速度要快得多。
MCVE 方面的代码是:
*** Open original dataset and work with it ***
webuse voter, clear
label list // shows two variables with value labels (candidat and inc)
drop candidat inc
label drop candidat inc2 // we drop all value labels
save final.dta
*** Create temporary dataset ***
use pop frac candidat using http://www.stata-press.com/data/r14/voter, clear // this is key. Only load needed variables!
*** Merge temporary dataset with original one ***
merge 1:1 pop frac using final.dta, nogen
label list // we only have the "candidat" label! Success!
save final.dta, replace