【问题标题】:an error in Recurrent Neural Network(LSTM) for tweet classification in PythonPython 中用于推文分类的循环神经网络 (LSTM) 中的错误
【发布时间】:2021-11-11 00:17:14
【问题描述】:

我正在尝试改进 LSTM 的结果。在我的部分项目中,我为 RNN 做了以下工作:

以下是用于训练模型的快速方法:

    def threshold_search(y_true, y_proba, average = None):
        best_threshold = 0
        best_score = 0
        for threshold in [i * 0.01 for i in range(100)]:
            score = f1_score(y_true=y_true, y_pred=y_proba > threshold, average=average)
            if score > best_score:
                best_threshold = threshold
                best_score = score
        search_result = {'threshold': best_threshold, 'f1': best_score}
        return search_result
    def train(model, 
              X_train, y_train, X_test, y_test, 
              checkpoint_path='model.hdf5', 
              epcohs = 25, 
              batch_size = DEFAULT_BATCH_SIZE, 
              class_weights = None, 
              fit_verbose=2,
              print_summary = True
             ):
        m = model()
        if print_summary:
            print(m.summary())
        m.fit(
            X_train, 
            y_train, 
            #this is bad practice using test data for validation, in a real case would use a seperate validation set
            validation_data=(X_test, y_test),
            epochs=epcohs, 
            batch_size=batch_size,
            class_weight=class_weights,
             #saves the most accurate model, usually you would save the one with the lowest loss
            callbacks= [
                ModelCheckpoint(checkpoint_path, monitor='val_acc', verbose=1, save_best_only=True),
                EarlyStopping(patience = 2)
            ],
            verbose=fit_verbose
        ) 
        print("\n\n****************************\n\n")
        print('Loading Best Model...')
        m.load_weights(checkpoint_path)
        predictions = m.predict(X_test, verbose=1)
        print('Validation Loss:', log_loss(y_test, predictions))
        print('Test Accuracy', (predictions.argmax(axis = 1) == y_test.argmax(axis = 1)).mean())
        print('F1 Score:', f1_score(y_test.argmax(axis = 1), predictions.argmax(axis = 1), average='weighted'))
        plot_confusion_matrix(y_test.argmax(axis = 1), predictions.argmax(axis = 1), classes=encoder.classes_)
        plt.show()    
        return m #returns best performing model

然后我使用了 LSTM 的简单实现。其中图层如下:

  • 嵌入:词向量矩阵,其中每个向量存储 “意义”这个词。这些可以即时训练或通过现有的 预训练向量。
  • LSTM:允许“构建” 随时间变化的状态
  • Dense(64):前馈神经网络用于 解释 LSTM 输出
  • Dense(3):这是模型的输出, 每个类对应3个节点。 softmax 输出将确保 每个输出的值总和 = 1.0。
def model_1():
    model = Sequential()
    model.add(Embedding(input_dim = (len(tokenizer.word_counts) + 1), output_dim = 128, input_length = MAX_SEQ_LEN))
    model.add(LSTM(128))
    model.add(Dense(64, activation='relu'))
    model.add(Dense(3, activation='softmax'))
    model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
    return model

m1 = train(model_1, 
           train_text_vec,
           y_train,
           test_text_vec,
           y_test,
           checkpoint_path='model_1.h5',
           class_weights= model.any(cws))

但我得到以下输出和错误:

Screenshot of the error

正如你在截图中看到的,错误是:

ValueError:具有多个元素的数组的真值是 模糊的。使用 a.any() 或 a.all()

你能帮我解决这个错误吗?

【问题讨论】:

  • 请将完整错误跟踪(单击“4 帧”将它们全部展开)发布为文本

标签: python machine-learning deep-learning nlp lstm


【解决方案1】:

基于 Keras 文档和 this questionclass_weights 需要一个字典,将整数类索引映射到浮点数表示的权重。

我不确定model.any(cws) 行应该做什么,但通常.any() 函数返回一个布尔值或布尔值数组。由于 class_weights 需要一个 dict,它会恐慌并抛出一个 ValueError。

我的猜测是,您将模型权重(构成模型的数字)与类别权重(您尝试预测的事物的相对重要性)混淆了。如果是这种情况,将model_weights 设置为默认值应该可以解决您的问题。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2016-09-07
    • 2017-08-11
    • 1970-01-01
    • 1970-01-01
    • 2018-10-05
    • 1970-01-01
    • 2017-02-19
    相关资源
    最近更新 更多