【发布时间】:2022-02-15 21:34:56
【问题描述】:
X.shape #output is => (2555904, 1024, 2)
X[0] #Output is => array([[ 0.0420274 , 0.23476323], [-0.2728826 , 0.40513492], [-0.26707262, 0.22749889], ..., [-0.7055947 , -0.28693035], [-0.41157472, 0.66826206], [ 0.06487698, 0.6358149 ]], dtype=float32)
total = len(X)
n_train = int(0.8*total) #80% samples in the training dataset 20% in testing
n_test = int(0.2*total)
train_idx = np.random.choice(range(0, total), size=n_train, replace=False) # Randomly selecting 80% of data from total dataset
test_idx = list(set(range(0, total)) - set(train_idx))
train_idx.sort()
test_idx.sort()
X_train = X[train_idx]
X_test = X[test_idx]
我被困在这段代码的最后两行,即 X_train 和 X_test 部分。运行这部分代码需要很多时间。还有另一种方法可以做到这一点吗?我想要的只是将 X 数据分成 80 20 个比率。欢迎提出任何建议。
我使用的数据集是 RadioML2018.01A。
同样的链接是:https://www.kaggle.com/pinxau1000/radioml2018-01a-get-started/data
我认为主要问题是数据的大小,如何克服它并隔离数据?
【问题讨论】:
-
没有像这样操作 20 GB 内存的“即时方式”。你更好的选择是只加载必要的部分(也就是一批数据)