【问题标题】:With working code: simple neural network (`y = 2 * x`)使用工作代码:简单的神经网络(`y = 2 * x`)
【发布时间】:2021-11-03 23:17:10
【问题描述】:

我是神经网络和机器学习的初学者(因此我希望模型学习简单的y = 2 * x)。我先解释一下我所知道的,最后是最小的工作示例。

生成器函数仅输出 (x, y) 与 y = 2 * x 对,并以 batch_size 为一组进行批处理:

generator = data_generator(batch_size=3)

print(next(generator))
# (
#    array([[0], [1], [2]]),
#    array([[0], [2], [4]]),
# )

print(next(generator))
# (
#    array([[3], [4], [5]]),
#    array([[6], [8], [10]]),
# )

print(next(generator))
# (
#    array([ [6],  [7],  [8]]),
#    array([[12], [14], [16]]),
# )

由于这是一个线性函数,模型只需要单层就可以学习y = 2 * x:

  • 输入
  • 乘数层:tf.keras.layers.Dense(1)(其中权重应为2,来自y = 2 * x)
  • 输出

运行以下代码,只需要pip install tensorflow==2.6.0

最小的工作示例(运行,但准确度始终接近 0):

import random
import tensorflow as tf
import numpy as np


def data_generator(batch_size):
    start = 0
    end = start + batch_size

    while True:
        x = np.arange(start, end)
        y = 2 * np.arange(start, end)

        x = np.expand_dims(x, axis=-1)
        y = np.expand_dims(y, axis=-1)

        yield x, y

        start += batch_size
        end = start + batch_size

        # Randomly divide so that we get repeated data points
        # if random.random() < 0.5:
        #     start //= 2


model = tf.keras.Sequential([
    tf.keras.layers.InputLayer(input_shape=(1,)),
    tf.keras.layers.Dense(1),
])

model.compile(
    loss=tf.keras.losses.BinaryCrossentropy(),
    optimizer=tf.keras.optimizers.Adam(),
    metrics=['accuracy'],
)

batch_size = 32

generator_training = data_generator(batch_size)
generator_testing = data_generator(batch_size)

model.fit(
    x=generator_training,
    validation_data=generator_testing,
)

我尝试了以下方法来提高准确性,但无济于事:

  • 使batch_size 为generator_testing 小于batch_size 为generator_training 以便测试通常在已经训练过的数据上进行
  • 使用if random.random() &lt; 0.5: start //= 2,这样训练会得到一些数据点不止一次
  • 添加更多Dense层
  • 改变每个Dense层中的神经元数量
  • 将所有 x 和 y 值除以一个大数 (1000000),使这些值介于 0 和 1 之间(至少在开始时)

为什么训练给出的准确率是10e-4 或更低,我应该怎么做才能让神经网络学习y = 2 * x?

【问题讨论】:

  • 这是一个回归问题。精度对于回归问题没有意义。我建议您先研究 ML 理论,然后确定合适的损失函数和度量标准。
  • 如果你不使用适当的损失函数,模型将无法训练,这里的二元交叉熵没有意义,你应该使用均方误差。

标签: python machine-learning keras neural-network tf.keras


【解决方案1】:
  • 对于这样的任务,不需要生成器,您可以直接传递 numpy 数组
  • 生成器也是一个无限循环,因此模型不断更新 数据并保持在同一时期
  • 模型也太简单了。
  • 要改进模型,它必须根据之前遇到的训练数据更新其权重,但由于生成器不断生成新数据并且它停留在第一个 epoch,因此没有任何改进。


下面的代码有

  • 使用序列 API 的数据生成器
  • 使用了不同的模型
  • 向层添加了激活函数
  • 使用均方误差损失函数

import random
import tensorflow as tf
import keras
import numpy as np
from keras.models import Sequential
from keras.layers import Dense
from keras import regularizers
from matplotlib import pyplot as plt

class DataGenerator(tf.keras.utils.Sequence):
    'Generates data for Keras'
    def __init__(self, x, y, batch_size=32):
        'Initialization'
        self.x = x
        self.y = y
        self.batch_size = batch_size

    def __len__(self):
        'Denotes the number of batches per epoch'
        return int(np.floor(len(self.x) / self.batch_size))

    def __getitem__(self, index):
        'Generate one batch of data'
        data = self.x[index*self.batch_size:(index+1)*self.batch_size]
        labels = self.y[index*self.batch_size:(index+1)*self.batch_size]

        return data, labels

model = Sequential()
model.add(Dense(8, activation='relu', kernel_regularizer=regularizers.l2(0.001), input_shape = (1,)))
model.add(Dense(8, activation='relu', kernel_regularizer=regularizers.l2(0.001)))
model.add(Dense(1))

optimizer = tf.keras.optimizers.Adam()
model.compile(optimizer=optimizer,loss='mse', metrics=['loss'])

# Print the model summary
model.summary()

# train data
x = np.arange(0, batch_size * 100)
y = 2 * np.arange(0, batch_size * 100)
train = DataGenerator(x, y, batch_size)

n_epochs = 100
batch_size = 32

history = model.fit(train, batch_size=batch_size, epochs=n_epochs,)

# plot loss during training
plt.figure()
plt.title('Loss / Mean Squared Error')
plt.plot(history.history['loss'], label='train')
plt.legend()
plt.show()

plt.figure(figsize=(20, 5))

plt.subplot(1,3,1)
plt.xlabel('x')
plt.ylabel('y')
plt.plot(x, y,)
plt.title('Original data')

y_predict_train = tf.math.ceil(model.predict(x))
plt.subplot(1,3,2)
plt.plot(x, y_predict_train,)
plt.xlabel('x')
plt.ylabel('prediction')
plt.title('Predictions')

plt.show()

# test data
x_test = np.arange(batch_size * 1000, batch_size * 1500)
y_test = 2 * np.arange(batch_size * 1000, batch_size * 1500)

# floor gives the lower end of the floating point number ex. 4.560 will give 4
# you could also use ceil function ex 4.560 will give 5
y_predict_test = tf.math.floor(model.predict(x_test))

plt.figure(figsize=(20, 5))

plt.subplot(1,3,1)
plt.xlabel('x')
plt.ylabel('y')
plt.plot(x_test, y_test,)
plt.title('Test data')

plt.subplot(1,3,2)
plt.plot(x_test, y_predict_test,)
plt.xlabel('x')
plt.ylabel('prediction')
plt.title('Predictions')
plt.show()

q = model.predict( np.array( [10,200, 360, 1000] ) )
print(q)
print(tf.math.floor(q))

【讨论】:

  • 请注意,这里唯一真正重要的是损失函数(可能还有数据生成,没看过)。由于询问者只是试图近似y = 2x,因此不需要隐藏层。
  • @xdurch0 我已经尝试了这个答案提供的代码并更新了我的代码以使用 MSE,你说得对,只需更新损失函数就可以了
猜你喜欢
  • 2011-04-16
  • 2019-01-23
  • 2016-08-08
  • 2022-01-08
  • 2015-10-04
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-03-02
相关资源
最近更新 更多