【问题标题】:Training a neural network to add训练神经网络以添加
【发布时间】:2011-05-11 10:11:03
【问题描述】:

我需要训练一个网络来相乘或相加 2 个输入,但对于 20000 之后的所有点,它似乎都不能很好地近似 迭代。更具体地说,我在整个数据集上对其进行训练,并且在最后一点上非常接近,但似乎 就像第一个端点并没有变得更好一样。我将数据标准化,使其介于 -0.8 和 0.8 之间。这 网络本身由 2 个输入、3 个隐藏神经元和 1 个输出神经元组成。我还将网络的学习率设置为 0.25, 并用作学习函数 tanh(x)。

对于数据集中最后训练的点,它非常近似,但对于第一个点,它看起来很像 不能很好地近似。我想知道它是什么,它不能很好地调整它,无论是我使用的拓扑,还是 还有什么?

还有多少神经元适合这个网络的隐藏层?

【问题讨论】:

  • 据我所知,神经元有二进制输出,它“触发”与否。如果默认输出为 1 或 0,您打算如何获得诸如加法或乘法之类的输出?
  • @dStulle 不,这只是您所说的一种神经元(即使是常见的)。

标签: neural-network


【解决方案1】:

由权重={1,1}、偏差=0 和线性激活函数的单个神经元组成的网络执行两个输入数字的相加。

乘法可能更难。以下是网络可以使用的两种方法:

  1. 将其中一个数字转换为数字(例如二进制)并像在小学时那样执行乘法运算。 a*b = a*(b0*2^0 + b1*2^1 + ... + bk*2^k) = a*b0*2^0 + a*b1*2^1 + ... + a*bk*2^k。这种方法很简单,但需要与输入 b 的长度(对数)成比例的可变数量的神经元。
  2. 取输入的对数,将它们相加并对结果取幂。 a*b = exp(ln(a) + ln(b)) 这个网络可以处理任何长度的数字,只要它能够很好地逼近对数和指数。

【讨论】:

    【解决方案2】:

    可能为时已晚,但一个简单的解决方案是使用 RNN (Recurrent Neural Network)。

    将您的数字转换为数字后,您的 NN 将从左到右的数字序列中取几个数字。

    RNN 必须循环其输出之一,以便它可以自动理解有一个数字要进位(如果和为 2,则写入 0 并进位 1)。

    要对其进行训练,您需要为其提供由两位数字组成的输入(一个来自第一个数字,第二个来自第二个数字)和所需的输出。 RNN 最终会找到求和的方法。

    请注意,此 RNN 只需要知道 以下 8 种情况即可学习如何对两个数字求和:

    • 1 + 1, 0 + 0, 1 + 0, 0 + 1 带进位
    • 1 + 1, 0 + 0, 1 + 0, 0 + 1 不带进位

    【讨论】:

      【解决方案3】:

      如果你想保持神经(链接有权重,神经元通过权重计算输入的总和,并根据总和的 sigmoid 回答 0 或 1,然后使用梯度的反向传播),那么您应该将隐藏层的神经元视为分类器。他们定义了一条线,将输入空间分成几类:一类对应于神经元响应 1 的部分,另一类对应于它响应 0 的部分。隐藏层的第二个神经元将定义另一个分隔,依此类推。输出神经元通过调整其输出的权重来组合隐藏层的输出,使其与您在学习期间呈现的输出相对应。
      因此,单个神经元会将输入空间分为 2 类(可能对应于根据学习数据库的添加)。两个神经元将能够定义 4 个类。三个神经元 8 类等。将隐藏神经元的输出视为 2 的幂:h1*2^0 + h2*2^1+...+hn*2^n,其中hi 是隐藏神经元i 的输出。注意:您将需要 n 个输出神经元。这回答了有关要使用的隐藏神经元数量的问题。
      但是 NN 不计算加法。它将其视为基于所学知识的分类问题。对于超出其学习基础的值,它将永远无法生成正确的答案。在学习阶段,它会调整权重以放置分隔符(2D 中的线),从而产生正确的答案。如果您的输入在[0,10] 中,它将学习为在[0,10]^2 中添加值生成正确答案,但永远不会为12 + 11 给出一个好的答案。
      如果您的最后一个值被很好地学习并且第一个被遗忘,请尝试降低学习率:最后一个示例的权重(取决于梯度)的修改可能会覆盖第一个(如果您使用的是随机反向传播)。确保您的学习基础是公平的。你也可以更频繁地展示那些学得不好的例子。并尝试许多学习率值,直到找到一个好的值。

      【讨论】:

      • 即NN 是为了解决分类问题,而加法不是分类问题 - 这是您在答案中暗示的吗?
      • 是的。在一般的情况下。它可以被视为一组受限数字的分类 pb(例如 {0,1} ;-)),但在一般情况下,它不是。
      【解决方案4】:

      我也试图这样做。训练了 2、3、4 位加法并能够达到 97% 的准确率。您可以使用其中一种神经网络类型来实现,

      Sequence to Sequence Learning with Neural Networks

      keras 的 Juypter Notebook 示例程序可在以下链接获得,

      https://github.com/keras-team/keras/blob/master/examples/addition_rnn.py

      希望对你有帮助。

      附上代码供参考。

      from __future__ import print_function
      from keras.models import Sequential
      from keras import layers
      import numpy as np
      from six.moves import range
      
      
      class CharacterTable(object):
          """Given a set of characters:
          + Encode them to a one hot integer representation
          + Decode the one hot integer representation to their character output
          + Decode a vector of probabilities to their character output
          """
          def __init__(self, chars):
              """Initialize character table.
              # Arguments
                  chars: Characters that can appear in the input.
              """
              self.chars = sorted(set(chars))
              self.char_indices = dict((c, i) for i, c in enumerate(self.chars))
              self.indices_char = dict((i, c) for i, c in enumerate(self.chars))
      
          def encode(self, C, num_rows):
              """One hot encode given string C.
              # Arguments
                  num_rows: Number of rows in the returned one hot encoding. This is
                      used to keep the # of rows for each data the same.
              """
              x = np.zeros((num_rows, len(self.chars)))
              for i, c in enumerate(C):
                  x[i, self.char_indices[c]] = 1
              return x
      
          def decode(self, x, calc_argmax=True):
              if calc_argmax:
                  x = x.argmax(axis=-1)
              return ''.join(self.indices_char[x] for x in x)
      
      
      class colors:
          ok = '\033[92m'
          fail = '\033[91m'
          close = '\033[0m'
      
      # Parameters for the model and dataset.
      TRAINING_SIZE = 50000
      DIGITS = 3
      INVERT = True
      
      # Maximum length of input is 'int + int' (e.g., '345+678'). Maximum length of
      # int is DIGITS.
      MAXLEN = DIGITS + 1 + DIGITS
      
      # All the numbers, plus sign and space for padding.
      chars = '0123456789+ '
      ctable = CharacterTable(chars)
      
      questions = []
      expected = []
      seen = set()
      print('Generating data...')
      while len(questions) < TRAINING_SIZE:
          f = lambda: int(''.join(np.random.choice(list('0123456789'))
                          for i in range(np.random.randint(1, DIGITS + 1))))
          a, b = f(), f()
          # Skip any addition questions we've already seen
          # Also skip any such that x+Y == Y+x (hence the sorting).
          key = tuple(sorted((a, b)))
          if key in seen:
              continue
          seen.add(key)
          # Pad the data with spaces such that it is always MAXLEN.
          q = '{}+{}'.format(a, b)
          query = q + ' ' * (MAXLEN - len(q))
          ans = str(a + b)
          # Answers can be of maximum size DIGITS + 1.
          ans += ' ' * (DIGITS + 1 - len(ans))
          if INVERT:
              # Reverse the query, e.g., '12+345  ' becomes '  543+21'. (Note the
              # space used for padding.)
              query = query[::-1]
          questions.append(query)
          expected.append(ans)
      print('Total addition questions:', len(questions))
      
      print('Vectorization...')
      x = np.zeros((len(questions), MAXLEN, len(chars)), dtype=np.bool)
      y = np.zeros((len(questions), DIGITS + 1, len(chars)), dtype=np.bool)
      for i, sentence in enumerate(questions):
          x[i] = ctable.encode(sentence, MAXLEN)
      for i, sentence in enumerate(expected):
          y[i] = ctable.encode(sentence, DIGITS + 1)
      
      # Shuffle (x, y) in unison as the later parts of x will almost all be larger
      # digits.
      indices = np.arange(len(y))
      np.random.shuffle(indices)
      x = x[indices]
      y = y[indices]
      
      # Explicitly set apart 10% for validation data that we never train over.
      split_at = len(x) - len(x) // 10
      (x_train, x_val) = x[:split_at], x[split_at:]
      (y_train, y_val) = y[:split_at], y[split_at:]
      
      print('Training Data:')
      print(x_train.shape)
      print(y_train.shape)
      
      print('Validation Data:')
      print(x_val.shape)
      print(y_val.shape)
      
      # Try replacing GRU, or SimpleRNN.
      RNN = layers.LSTM
      HIDDEN_SIZE = 128
      BATCH_SIZE = 128
      LAYERS = 1
      
      print('Build model...')
      model = Sequential()
      # "Encode" the input sequence using an RNN, producing an output of HIDDEN_SIZE.
      # Note: In a situation where your input sequences have a variable length,
      # use input_shape=(None, num_feature).
      model.add(RNN(HIDDEN_SIZE, input_shape=(MAXLEN, len(chars))))
      # As the decoder RNN's input, repeatedly provide with the last hidden state of
      # RNN for each time step. Repeat 'DIGITS + 1' times as that's the maximum
      # length of output, e.g., when DIGITS=3, max output is 999+999=1998.
      model.add(layers.RepeatVector(DIGITS + 1))
      # The decoder RNN could be multiple layers stacked or a single layer.
      for _ in range(LAYERS):
          # By setting return_sequences to True, return not only the last output but
          # all the outputs so far in the form of (num_samples, timesteps,
          # output_dim). This is necessary as TimeDistributed in the below expects
          # the first dimension to be the timesteps.
          model.add(RNN(HIDDEN_SIZE, return_sequences=True))
      
      # Apply a dense layer to the every temporal slice of an input. For each of step
      # of the output sequence, decide which character should be chosen.
      model.add(layers.TimeDistributed(layers.Dense(len(chars))))
      model.add(layers.Activation('softmax'))
      model.compile(loss='categorical_crossentropy',
                    optimizer='adam',
                    metrics=['accuracy'])
      model.summary()
      
      # Train the model each generation and show predictions against the validation
      # dataset.
      for iteration in range(1, 200):
          print()
          print('-' * 50)
          print('Iteration', iteration)
          model.fit(x_train, y_train,
                    batch_size=BATCH_SIZE,
                    epochs=1,
                    validation_data=(x_val, y_val))
          # Select 10 samples from the validation set at random so we can visualize
          # errors.
          for i in range(10):
              ind = np.random.randint(0, len(x_val))
              rowx, rowy = x_val[np.array([ind])], y_val[np.array([ind])]
              preds = model.predict_classes(rowx, verbose=0)
              q = ctable.decode(rowx[0])
              correct = ctable.decode(rowy[0])
              guess = ctable.decode(preds[0], calc_argmax=False)
              print('Q', q[::-1] if INVERT else q, end=' ')
              print('T', correct, end=' ')
              if correct == guess:
                  print(colors.ok + '☑' + colors.close, end=' ')
              else:
                  print(colors.fail + '☒' + colors.close, end=' ')
              print(guess)
      

      【讨论】:

        【解决方案5】:

        想想如果你用 x 的线性函数替换 tanh(x) 阈值函数会发生什么 - 称之为 a.x - 并将 a 视为每个神经元中唯一的学习参数。这实际上就是您的网络将要优化的目标;这是tanh 函数过零的近似值。

        现在,当您将这种线性类型的神经元分层时会发生什么?当脉冲从输入到输出时,您每个神经元的输出。您正在尝试使用一组乘法来近似加法。正如他们所说,这不会计算。

        【讨论】:

        • 我认为你的数学有问题。如果没有非线性激活函数,多层 ANN 会坍缩成单层,而单层只是将输入向量乘以权重矩阵。每个输出只是输入的加权 sum。具有1 权重和 0 偏差的单个神经元/感知器将两个输入相加。
        猜你喜欢
        • 2011-04-07
        • 1970-01-01
        • 2017-05-28
        • 2010-11-20
        • 2019-09-15
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        相关资源
        最近更新 更多