【问题标题】:Is there a faster or better way to segregate dataset into 80 20 ratio in python?有没有更快或更好的方法在 python 中将数据集分成 80 20 比率?
【发布时间】:2022-02-15 21:34:56
【问题描述】:
X.shape  #output is => (2555904, 1024, 2)
X[0] #Output is => array([[ 0.0420274 , 0.23476323], [-0.2728826 , 0.40513492], [-0.26707262, 0.22749889], ..., [-0.7055947 , -0.28693035], [-0.41157472, 0.66826206], [ 0.06487698, 0.6358149 ]], dtype=float32)
total = len(X)
n_train = int(0.8*total) #80% samples in the training dataset 20% in testing
n_test = int(0.2*total)
train_idx = np.random.choice(range(0, total), size=n_train, replace=False) # Randomly selecting 80% of data from total dataset
test_idx = list(set(range(0, total)) - set(train_idx))
train_idx.sort()
test_idx.sort()

X_train = X[train_idx]
X_test = X[test_idx]

我被困在这段代码的最后两行,即 X_train 和 X_test 部分。运行这部分代码需要很多时间。还有另一种方法可以做到这一点吗?我想要的只是将 X 数据分成 80 20 个比率。欢迎提出任何建议。

我使用的数据集是 RadioML2018.01A。

同样的链接是:https://www.kaggle.com/pinxau1000/radioml2018-01a-get-started/data

我认为主要问题是数据的大小,如何克服它并隔离数据?

【问题讨论】:

标签: python pandas


【解决方案1】:

您可以使用 sklearn.model_selection.train_test_split。 查看official documentation 的示例。首先必须将数据拆分为目标变量和解释变量。

【讨论】:

    猜你喜欢
    • 2017-08-31
    • 2021-04-17
    • 1970-01-01
    • 2016-02-04
    • 2011-06-28
    • 1970-01-01
    • 2019-08-13
    • 1970-01-01
    • 2013-01-28
    相关资源
    最近更新 更多