【问题标题】:Python Backpropagation: Gradient becomes increasingly small for increasing batch sizePython 反向传播:随着批大小的增加,梯度变得越来越小
【发布时间】:2020-10-15 08:45:59
【问题描述】:

我正在训练以下维度的神经网络: 784(输入层) 45(隐藏层) 16(输出层)

对于数字和一些数学符号(0-9,+,-,*,/,[,])的分类,使用反向传播(随机梯度下降)

在mini-batch size的选择上我做了一些测试,发现了以下问题:

1.使用 20 个数据点的小批量大小,反向传播算法“似乎”可以工作,但即使在 50 多个 epoch 上对其进行训练后,准确度似乎只会波动并变得更糟(图 1)

2 . 使用 2000 个数据点的 mini-batch 大小,权重梯度非常小,以至于更新后它们并没有真正改变实际权重

下面我发布了我用来训练神经网络对象的类的相关代码。虽然不是所有的东西都是可见的,但名称是不言自明的。

一些相关数据:

  1. 训练数据集是约 200k 个数据点(28x28 numpy 数组和相应符号的元组)
  2. 验证数据集约为 50k 个数据点
  3. 该算法使用 MSE 作为成本函数
  4. 我正在使用以下公式的反向传播算法:

  1. 重要提示:为了使反向传播计算更有效,我将计算作为批处理张量操作执行,其中相应的梯度在张量中,其中第一个轴对应于数据集索引,其余的充当正常矩阵/向量。关于上一个问题的更多信息: Python: numpy.dot / numpy.tensordot for multidimensional arrays

示例(小批量:20)

在最后一层激活: 一个数据集是:(16x1) 批量反向传播是:(20x16x1)

最后一层的权重梯度: 对于一个数据集是:(16x45) 批量反向传播是:(20x16x45)

import numpy as np
import random as rd
import time
import matplotlib.pyplot as plt

class NeuralNetworkTrainer:
  def __init__(self, neuralNetwork,validator):
    self.network = neuralNetwork # Uses the neural network object which contains weights, bias' as a list of numpy arrays for each layer (the first element being 'None' to be consistent with indexes), as well as activations and layer sizes
    self.eta = 0
    self.dataSet = [] # Another class loads the dataset here (tuples of inputs, 2D-numpy arrays of the image and outputs, the corresponding symbol)
   
    self.initializeWeightsBias()
    self.validator = validator # validator object
    self.validationAccuracy = [] # list of accuracies per epoch
    
  def initializeWeightsBias(self): #gradients initialization
    self.gradientToBias = [None]*len(self.network.layers)
    self.gradientToWeights = [None]*len(self.network.layers)

  def train(self,epochs,miniBatchSize,eta): #train algorithm
    self.eta = eta
    for i in range(0,epochs):
      self.shuffleData()
      for j in range(0,len(self.dataSet)//miniBatchSize):
        self.batchBackPropagation(self.createMiniBatch(miniBatchSize,j))
        self.update()

      correctOutputs, dataSetLength = self.validator.validate()
      self.validationAccuracy.append(round(correctOutputs/dataSetLength,4))
    
    return self.network

# ***************************
# BACKPROPAGATION ALGORITHM

  def batchBackPropagation(self,inputOutputBatch):
    self.initializeWeightsBias()

    activations = [None]*len(self.network.activations)
    for i in range(0,len(activations)): #Initialize activations
      activations[i] = np.empty((len(inputOutputBatch),self.network.activations[i].shape[0],self.network.activations[i].shape[1]))
    
    output = np.empty((len(inputOutputBatch),self.network.activations[-1].shape[0],self.network.activations[-1].shape[1])) #correct formatting of output vector out of the symbol (vector with 0's and a 1 in the corresponding output)
    for i in range(0,len(inputOutputBatch)):
      inputVector, outputVector = self.vectorizeInputOuput(inputOutputBatch[i])
      self.network.loadInput(inputVector)
      self.network.activate() #feedforward of input through the network with current weights/bias
      output[i] = outputVector
      for l in range(1,len(activations)): #creation of activation tensor as explained before
        activations[l][i] = self.network.activations[l]
    
    self.gradientToBias[-1] =(activations[-1]-output)*(activations[-1]-np.square(activations[-1])) #calculation of gradientBias for last layer for all the minibatches as a 3D tensor calculation (see algorithm image)
    for i in range(2,len(self.network.layers)):
      self.gradientToBias[-i] = np.tensordot(self.gradientToBias[-i+1],self.network.weights[-i+1],axes= ((1),(0))).transpose(0,2,1)*(activations[-i]-np.square(activations[-i])) #calculation of the rest of the gradientToBias for the rest of the layers as a 3D tensor calculation the first index being the index of the dataset in that minibatch (according to algorithm image)
    for i in range(1,len(self.network.layers)):
      self.gradientToWeights[i] = np.einsum('ijk,ilm->ijl',self.gradientToBias[i],activations[i-1])
    return self.network # analogous 3D tensor calculation of gradientToWeights for each dataset in the minibatch inside every layer of the 3D tensor

# *****************************

  def update(self): #reduction of gradients of each dataset to one final gradient to each parameter by summing over axis=0)
    for i in range(1,len(self.network.layers)):
      self.network.weights[i] -= self.eta*np.sum(self.gradientToWeights[i],axis =0)
      self.network.bias[i] -= self.eta*np.sum(self.gradientToBias[i], axis = 0)
    return self.network

  def shuffleData(self): #self explanatory
    rd.shuffle(self.dataSet)
    return self.network 

  def createMiniBatch(self, miniBatchSize, index): #self explanatory
    return self.dataSet[index*miniBatchSize:(index+1)*miniBatchSize] 

  def mapOutputToVector(self,output): #self explanatory
      outputVector = np.zeros((len(self.network.outputMap),1))
      outputVector[self.network.outputMap.index(output)] = 1
      return outputVector

  def vectorizeInputOuput(self,inputOutputData): #selfexplanatory
    return inputOutputData.input.flatten().reshape((-1,1)), self.mapOutputToVector(inputOutputData.output)
  

非常感谢任何帮助!

【问题讨论】:

  • 要尝试的一件事是降低隐藏层。对于这个简单的问题,您的架构可能太大了。此外,还要在混合中添加一些 dropout,以帮助使模型更加通用。这看起来像是在学习太具体的特征,这就是它波动的原因。 2000 通常是太大的批量尝试这个​​范围 [16:526] 这些是一些文献中的最佳选择。对于您的小型数据集,我会坚持使用 32 或 64。
  • 只有一个隐藏层。你的意思是缩小它的大小?谢谢!
  • 我看我读错了'45(隐藏层)'我会试试这个 784->64->32->16 这个设置对于我构建的 pytorch 模型很有效
  • 我见过的使用 MNIST 库来识别手写数字的示例通常带有一个隐藏层。您确定放置两层并将参数数量增加 ca. 50% 是不是只对另外 6 个字符过于复杂?
  • @MichelH。使用 cross-entropy loss 作为成本函数,MSE 不能很好地解决分类问题。对于激活函数,隐藏层使用 relu 或 tanh,输出层使用 softmax。还可以使用 for 循环并尝试使用不同的 eta 值进行训练。

标签: python machine-learning neural-network gradient-descent backpropagation


【解决方案1】:

首先,过大的小批量通常会导致准确性降低。

您在第一张图中面临的问题是过度拟合,因此您需要减少 epoch 的数量。

至于第二个图,您有 200,000 个样本和 2000 的批量大小,那么 epoch 应该包含 200,000 / 2000 = 100 个步骤,这被认为是每个 epoch 梯度中的少量步骤

一般来说,您需要为时期数和批量大小选择正确的数字以获得最佳结果。也许 1000 步在 epoch 中,不要训练太多 epoch 以避免过度拟合

【讨论】:

  • 非常感谢!但我并不完全相信你的回答。为什么较大的小批量会降低准确性?由于内存问题,我们首先创建了小批量。在没有内存限制的情况下,我们将“选择”小批量作为完整的训练集,然后在各个时期进行迭代。此外,在 20 个 epoch 之后,我无法使用具有大约 35k 参数的神经网络过度拟合 200k 数据集。最大的问题是为什么权重梯度变得如此之小?我在矩阵初始化方面做得对吗?我是否错误地操作了 num 数组?
  • 这个link讨论mini-batch size的效果看看。至于代码本身,最​​简单的做法是使用具有相同超参数的 keras 或 tensorflow 构建另一个神经网络并比较结果。 (顺便说一句,我在代码中看不到任何错误)。但我仍然认为,数据集的分布可能会过度拟合,而不仅仅是样本的数量
猜你喜欢
  • 2021-01-23
  • 2017-02-08
  • 2015-04-13
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多