【问题标题】:How can I fix the train set and test set in python?如何在 python 中修复训练集和测试集?
【发布时间】:2019-02-24 17:03:27
【问题描述】:

假设我有一个包含 1000 行的数据集。我想把它分成训练集和测试集。我想将前 800 行拆分为训练集,然后将 200 行拆分为测试集。有可能吗?

我的训练和测试拆分的python测试代码是这样的:

from sklearn.cross_validation import train_test_split

xtrain, xtest, ytrain, ytest = train_test_split(x, y, test_size=0.20)

【问题讨论】:

  • 如果我理解正确,你不想要 train_test_split 行为的洗牌,如果是这样,你是使用 numpy 还是 pandas,还是其他什么?
  • 熊猫。我只是想确保我的前 800 个数据(按顺序)将在火车部分,然后其余 200 个在测试部分。

标签: python-3.x jupyter-notebook


【解决方案1】:

有多种方法可以做到这一点,我将通过其中一些来运行。

切片是python中一种强大的方法,如果您只想要前800个副本并且您的数据框被命名为train用于输入特征和Y 可以使用的输出功能

X_train = train[0:800]
X_test = train[800:]
y_train = Y[0:800]
y_test = Y[800:]

Iloc 函数与数据帧相关联并与索引相关联,如果您的索引是数字,那么您可以使用

X_train = train.iloc[0:800]
X_test = train.iloc[800:]
y_train = Y.iloc[0:800]
y_test = Y.iloc[800:]

如果你只需要将数据分成两部分,你甚至可以使用df.head() 和df.tail() 来做到这一点,

X_train = train.head(800)
X_test = train.tail(200)
y_train = Y.head(800)
y_test = Y.tail(200)

还有其他方法可以做到这一点,我建议使用第一种方法,因为它在多种数据类型中很常见,并且如果您使用 numpy 数组也可以使用。要了解有关切片的更多信息,我建议您结帐。 Understanding slice notation 这里解释为一个列表,但它几乎适用于所有形式。

【讨论】:

    【解决方案2】:

    你想设置shuffle= False。

    from sklearn.cross_validation import train_test_split
    xtrain, xtest, ytrain, ytest = train_test_split(x, y, test_size=0.20, shuffle = False) 
    

    【讨论】:

      猜你喜欢
      • 2020-06-11
      • 2016-03-22
      • 1970-01-01
      • 2017-11-01
      • 2015-04-10
      • 2015-01-17
      • 2013-06-18
      • 2019-08-15
      • 2017-10-01
      相关资源
      最近更新 更多