【问题标题】:In Random under sampling, How can I define drop ratio?在随机抽样中,如何定义丢弃率?
【发布时间】:2021-03-07 19:00:41
【问题描述】:

当我使用欠采样代码时,但似乎将主要类降低到与主要类的数量与次要类的数量相同的比例。(50% vs 50%)

当有 30% 的 Minor 课程时,我想为 Major 课程赚取 70%。

我该如何处理这个问题,在 Major 和 Minor Class 之间设置权重的参数是什么?

sampler = RandomUnderSampler(ratio={1: 1000, 0: 65})
X_rs, y_rs = sastrong textmpler.fit_sample(X, y)
print('Random undersampling {}'.format(Counter(y_rs)))

【问题讨论】:

    标签: imblearn


    【解决方案1】:

    回答

    所以基本上,RandomUnderSampler(sampling_strategy = X) 使用的策略是少数类占多数类的 X%。因此,如果您选择 X=1,您将获得类似于 auto 的结果,这使得两个类 100% 平衡。

    现在,如果您选择 0.9,您将使 minority 类成为 majority 类的 90%。

    因此,如果您希望总集为 70%-30%,则需要做一些数学运算(笑话)。

    sampling_strategy = 0.5,因为我们使用的是比率,如果多数类是少数类的两倍,我们会得到 1/3 与 2/3,这大约是您想要的。

    TLDR

    sampling_strategy 采用比率,而您想要 30/70 的比率。因此,您必须通过 3/7 即 0.428 。

    【讨论】:

      猜你喜欢
      • 2018-07-11
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2017-06-28
      相关资源
      最近更新 更多