【问题标题】:Input arrays should have the same number of samples as target arrays输入数组应具有与目标数组相同数量的样本
【发布时间】:2018-04-15 20:03:10
【问题描述】:

我在训练网络时遇到错误:Input arrays should have the same number of samples as target arrays. Found 50000 input samples and 25000 target samples. x_train 形状 (50000, 200) y_train (25000,) x_test (50000, 200) y_test (25000,) 我该如何纠正错误? 我用 keras 编写的网络 从 csv 文件加载数据集

将字符串列表转换为整数列表

x_train = [[int(elt) for elt in sublist] for sublist in x_train]

x_test = [[int(elt) for elt in sublist] for sublist in x_test]
y_train = [int(elt) for elt in y_train]
y_test = [int(elt) for elt in y_test]


# convert integer reviews into word reviews
x_full = x_train + x_test
x_full_words = [[index_to_word[idx] for idx in rev if idx!=0] for rev in x_full]
all_words = [word for rev in x_full_words for word in rev]
if use_pretrained:

    # initialize word vectors
    word_vectors = Word2Vec(size=word_vector_dim, min_count=1)

    # create entries for the words in our vocabulary
    word_vectors.build_vocab(x_full_words)

    # sanity check
    assert (len(list(set(all_words))) == len(word_vectors.wv.vocab)), "3rd sanity check failed!"

    # fill entries with the pre-trained word vectors
    word_vectors.intersect_word2vec_format(path_to_pretrained_wv + 'GoogleNews-vectors-negative300.bin.gz', binary=True)

    print('pre-trained word vectors loaded')

    norms = [np.linalg.norm(word_vectors[word]) for word in
             list(word_vectors.wv.vocab)]  # in Python 2.7: word_vectors.wv.vocab.keys()
    idxs_zero_norms = [idx for idx, norm in enumerate(norms) if norm < 0.05]
    no_entry_words = [list(word_vectors.wv.vocab)[idx] for idx in idxs_zero_norms]
    print('# of vocab words w/o a Google News entry:', len(no_entry_words))

    # create numpy array of embeddings
    embeddings = np.zeros((max_features + 1, word_vector_dim))
    for word in list(word_vectors.wv.vocab):
        idx = word_to_index[word]
        # word_to_index is 1-based! the 0-th row, used for padding, stays at zero
        embeddings[idx,] = word_vectors[word]

    print('embeddings created')

else:
    print('not using pre-trained embeddings')

和模型

#model
my_input = Input(shape=(max_size,)) # we leave the 2nd argument of shape blank because the Embedding layer cannot accept an input_shape argument

if use_pretrained:
    embedding = Embedding(input_dim=embeddings.shape[0], # vocab size, including the 0-th word used for padding
                          output_dim=word_vector_dim,
                          weights=[embeddings], # we pass our pre-trained embeddings
                          input_length=max_size,
                          trainable=not do_static,
                          ) (my_input)
else:
    embedding = Embedding(input_dim=max_features + 1,
                          output_dim=word_vector_dim,
                          trainable=not do_static,
                          ) (my_input)

embedding_dropped = Dropout(drop_rate)(embedding)

# feature map size should be equal to max_size-filter_size+1
# tensor shape after conv layer should be (feature map size, nb_filters)
print('branch A:',nb_filters,'feature maps of size',max_size-filter_size_a+1)
print('branch B:',nb_filters,'feature maps of size',max_size-filter_size_b+1)

# A branch
conv_a = Conv1D(filters = nb_filters,
              kernel_size = filter_size_a,
              activation = 'relu',
              )(embedding_dropped)

pooled_conv_a = GlobalMaxPooling1D()(conv_a)

pooled_conv_dropped_a = Dropout(drop_rate)(pooled_conv_a)

# B branch
conv_b = Conv1D(filters = nb_filters,
              kernel_size = filter_size_b,
              activation = 'relu',
              )(embedding_dropped)

pooled_conv_b = GlobalMaxPooling1D()(conv_b)

pooled_conv_dropped_b = Dropout(drop_rate)(pooled_conv_b)

concat = Concatenate()([pooled_conv_dropped_a,pooled_conv_dropped_b])

concat_dropped = Dropout(drop_rate)(concat)

# we finally project onto a single unit output layer with sigmoid activation
prob = Dense(units = 1, # dimensionality of the output space
             activation = 'sigmoid',
             )(concat_dropped)

model = Model(my_input, prob)
model.layers[4].output_shape # dimensionality of document encodings (nb_filters*2)

【问题讨论】:

  • 您必须发布更多源代码,尤其是在加载 x 和 y 数据集的地方
  • 你需要 x 和 y 长度相同,它们需要一对一,一个 x 对应一个 y

标签: python neural-network keras


【解决方案1】:

根据您对数据的描述,我无法确切说明问题出在哪里,但每个输入样本 (x) 必须有一个目标 (y),这就是您收到错误 Input arrays should have the same number of samples as target arrays. Found 50000 input samples and 25000 target samples. 的原因。您的形状(以及 Keras 错误)显示您有 50,000 个实例和只有 25,000 个目标。

要么是 csv 数据出错,要么是导入或处理过程中的某个地方出错。

【讨论】:

    猜你喜欢
    • 2018-07-10
    • 1970-01-01
    • 2018-09-28
    • 2020-06-13
    • 1970-01-01
    • 1970-01-01
    • 2020-08-31
    • 2023-03-10
    • 2019-05-30
    相关资源
    最近更新 更多