【问题标题】:Error with input shape in Keras while training CBOW model训练 CBOW 模型时 Keras 中的输入形状出错
【发布时间】:2020-08-12 08:23:32
【问题描述】:

我正在为单词嵌入训练一个连续的单词模型,其中每个 one-hot 向量的形状是一个形状为 (V, 1) 的列向量。我正在使用生成器根据语料库生成训练示例和标签,但输入形状有错误。

(这里 V = 5778)

这是我的代码:

def windows(words, C):
    i = C
    while len(words) - i > C:
        center = words[i]
        context_words = words[i-C:i] + words[i+1:i+C+1]
        i += 1
        yield context_words, center

def one_hot_rep(word, word_to_index, V):
    vec = np.zeros((V, 1))
    vec[word_to_index[word]] = 1
    return vec

def context_to_one_hot(words, word_to_index, V):
    arr = [one_hot_rep(w, word_to_index, V) for w in words]
    return np.mean(arr, axis=0)
def get_training_examples(words, C, words_to_index, V):
    for context_words, center_word in windows(words, C):
        yield context_to_one_hot(context_words, words_to_index, V), one_hot_rep(center_word, words_to_index, V)
V = len(vocab)
N = 50

w2i, i2w = build_dict(vocab)

model = keras.models.Sequential([
    keras.layers.Flatten(input_shape=(V, )),
    keras.layers.Dense(units=N, activation='relu'),
    keras.layers.Dense(units=V, activation='softmax')
])

model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])

model.fit_generator(get_training_examples(data, 2, w2i, V), epochs=5, steps_per_epoch=20)

【问题讨论】:

    标签: numpy keras nlp word-embedding


    【解决方案1】:

    flatten layer 得到至少 3 维的 numpy 数组,但你给它 2 维

    【讨论】:

    • 您应该更改此代码中的 V:keras.layers.Flatten(input_shape=(V, ))
    【解决方案2】:

    我找出了导致错误的原因。该模型需要一个 input_shape = (None, V) ,其中 None 在训练开始时保存 Keras 的 batch_size 但我发送的是形状为 (1, V) 的数组,当作为批处理发送时会获得额外的第一维,例如 ( 128, 1, V) 正在发送,与预期的 input_shape 冲突。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2021-01-24
      • 2019-05-06
      • 2019-05-19
      • 2019-10-20
      • 1970-01-01
      • 1970-01-01
      • 2017-05-29
      • 2022-01-12
      相关资源
      最近更新 更多