【问题标题】:make prediction in new dataset在新数据集中进行预测
【发布时间】:2019-07-29 09:52:05
【问题描述】:

我建立了一个 keras 逻辑回归模型。我正在尝试找到一种方法,可以为我的模型提供新的数据集,并在我通过的新数据集中进行预测。我的新数据集将与我的模型形状相同

我的第二个问题是有没有办法提高我的模型的准确性,因为我的准确率是 69%,当我打印分类报告时,我在一类中得到了不好的准确率

X=new.drop('reassed',axis=1)
y=new['reassed'].astype(int)

拆分数据

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2, random_state = 0)


from sklearn.preprocessing import StandardScaler
sc = StandardScaler()
X_train = sc.fit_transform(X_train)
X_test = sc.transform(X_test)

# Initialising the ANN
classifier = Sequential()

# Adding the input layer and the first hidden layer
classifier.add(Dense(units = 27, kernel_initializer = 'uniform', activation = 'relu', input_dim = 6))

# Adding the second hidden layer
classifier.add(Dense(units = 27, kernel_initializer = 'uniform', activation = 'relu'))

# Adding the output layer
classifier.add(Dense(units = 1, kernel_initializer = 'uniform', activation = 'sigmoid'))

# Compiling the ANN
classifier.compile(optimizer = 'adam', loss = 'binary_crossentropy', metrics = ['accuracy'])`enter code here`

# Fitting the ANN to the Training set
classifier.fit(X_train, y_train, batch_size = 10, epochs = 20)


Epoch 1/20
16704/16704 [==============================] - 1s 76us/step - loss: 0.6159 - acc: 0.6959
Epoch 2/20
16704/16704 [==============================] - 1s 65us/step - loss: 0.6114 - acc: 0.6967
Epoch 3/20
16704/16704 [==============================] - 1s 65us/step - loss: 0.6110 - acc: 0.6964
Epoch 4/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6101 - acc: 0.6965
Epoch 5/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6091 - acc: 0.6961
Epoch 6/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6094 - acc: 0.6963
Epoch 7/20
16704/16704 [==============================] - 1s 68us/step - loss: 0.6086 - acc: 0.6967
Epoch 8/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6083 - acc: 0.6965
Epoch 9/20
16704/16704 [==============================] - 1s 65us/step - loss: 0.6081 - acc: 0.6964: 0s - loss: 0.6085 - acc: 
Epoch 10/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6082 - acc: 0.6971
Epoch 11/20
16704/16704 [==============================] - 1s 67us/step - loss: 0.6077 - acc: 0.6968
Epoch 12/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6073 - acc: 0.6971
Epoch 13/20
16704/16704 [==============================] - 1s 65us/step - loss: 0.6067 - acc: 0.6971
Epoch 14/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6070 - acc: 0.6965
Epoch 15/20
16704/16704 [==============================] - 1s 65us/step - loss: 0.6066 - acc: 0.6967: 0s - loss: 0.6053 - ac
Epoch 16/20
16704/16704 [==============================] - 1s 66us/step - loss: 0.6060 - acc: 0.6967
Epoch 17/20
16704/16704 [==============================] - 1s 67us/step - loss: 0.6061 - acc: 0.6968
Epoch 18/20
16704/16704 [==============================] - 1s 67us/step - loss: 0.6062 - acc: 0.6971
Epoch 19/20
16704/16704 [==============================] - 1s 69us/step - loss: 0.6057 - acc: 0.6968
Epoch 20/20
16704/16704 [==============================] - 1s 74us/step - loss: 0.6055 - acc: 0.6973

y_pred = classifier.predict(X_test)
y_pred = [ 1 if y>=0.5 else 0 for y in y_pred ]

print(classification_report(y_test, y_pred))

      precision    recall  f1-score   support

           0       0.71      1.00      0.83      2968
           1       0.33      0.00      0.01      1208

   micro avg       0.71      0.71      0.71      4176
   macro avg       0.52      0.50      0.42      4176
weighted avg       0.60      0.71      0.59      4176

我希望改进我的模型

我希望找到一种可以在新数据集中进行预测的方法

【问题讨论】:

  • 低精度类的数据样本量是多少?对于您关于新数据集的问题,请使其通过执行的预处理步骤,并在加载权重后使用新数据集简单地调用预测函数。
  • 我猜1类精度低的原因是数据不平衡。0类的计数是14590,1类的计数是6290。如果你有,请在keras预测功能中提供我谢谢
  • 精度低是由于数据不平衡造成的。对于新数据集,执行与 X_test 相同的代码。
  • 我说不需要任何特定的源代码。你可以使用相同的预测函数
  • 我会将此作为答案,您可以将其标记为已接受的答案,以便对未来的用户有用。

标签: python machine-learning keras deep-learning prediction


【解决方案1】:

对新数据集进行预测

  1. 加载与加载测试集相同的数据
  2. 应用在您的训练集上应用的所有每个处理步骤。
  3. 使用

    model.predict(X)

功能进行预测并继续您的后期处理。

这与使用测试集进行预测几乎相同。

【讨论】:

  • 如果答案对您有效,请勾选左侧的勾号,将其标记为已接受。谢谢。 :)
猜你喜欢
  • 2021-05-31
  • 2016-06-12
  • 2019-12-25
  • 2018-10-15
  • 2023-03-18
  • 2018-11-29
  • 2018-07-21
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多