【问题标题】:Custom metric for Keras model, using Tensorflow 2.1Keras 模型的自定义指标,使用 Tensorflow 2.1
【发布时间】:2021-03-26 14:31:02
【问题描述】:

我想添加一个自定义指标来使用 Keras 建模,我正在调试我的工作代码,但找不到执行所需操作的方法。

该问题可以描述为通过逻辑多项式回归进行的多分类。 我想实现的自定义指标是这样的:

(1/Number_of_Classes)*(TruePositivesClass1/TotalElementsClass1 + TruePositivesClass2/TotalElementsClass2 + ... + TruePositivesClassN/TotalElementsClassN)

其中 Number_of_Classes 必须从批处理中计算,例如 np.unique(y_true).count() 和 每个总和项目都类似于

len(np.where(y_true==class_i,1,0) == np.where(y_pred==class_i,1,0) )/np.where(y_true==class_i,1,0).sum()

就混淆矩阵而言(2个变量的最小形式)

        True    False
True     15      3
False    12      1

公式为0.5*(15)/(15+12) + 0.5*(1/(1+3))=0.4027

代码可能类似于

def custom_metric(y_true,y_pred):

    total_classes = Unique(y_true) #How calculate total unique elements?
    summation = 0
    for _ in unique_value_on_target:

        # calculates Number of y_predict that are _
        true_predics_of_class = Count(y_predict,_) 

        # calculates total number of items of class _ in batch y_true
        true_values = Count(y_true,_) 

        value = true_predicts/true_values
       summation + = value
    return summation

我的预处理数据是一个像 x=[v1,v2,v3,v4,...,vn] 这样的 numpy 数组,而我的 目标列是一个随机数组y=[1, 0, 1, 0, 1, 0, 0, 1 ,..., 0, 1]

然后,它们被转换为张量:

x_train = tf.convert_to_tensor(x)
y_train = tf.convert_to_tensor(tf.keras.utils.to_categorical(y))

然后,将它们转换为 tensorflow 数据集对象:

train_ds = tf.data.Dataset.zip((tf.data.Dataset.from_tensor_slices(x_train),
                                tf.data.Dataset.from_tensor_slices(y_train)))

稍后,我取了一个迭代器:

 train_itr = iter(
          train_ds.shuffle(len(y_train) * 5, reshuffle_each_iteration=True).batch(len(y_train)))

最后,我采用迭代器的一个元素并训练

x_train, y_train = train_itr.get_next()
model.fit(x=x_train, y=y_train, batch_size=batch_size, epochs=epochs,
          callbacks=[custom_callback], validation_data=test_itr.get_next())

因此,由于对象是数据集迭代器,因此我无法找到按我的意愿操作它们的函数,以获取所描述的自定义指标。

【问题讨论】:

    标签: python keras tensorflow2.x keras-metrics


    【解决方案1】:

    所以你想计算批量中多类的平均召回率,这是我使用numpy 和tensorflow 的示例代码:

    import tensorflow as tf
    import numpy as np
    
    y_t = np.array([[1, 0, 0, 0], [0, 1, 0, 0], [0, 1, 0, 0], [0, 0, 0, 1], [0, 0, 0, 1]], dtype=np.float32)
    y_p = np.array([[1, 0, 0, 0], [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 0, 1], [0, 0, 0, 1]], dtype=np.float32)
    
    def average_recall(y_true, y_pred):
        # Get indexes of both labels and predictions
        labels = np.argmax(y_true, axis=1)
        predictions = np.argmax(y_pred, axis=1)
        # Get confusion matrix from labels and predictions
        confusion_matrix = tf.math.confusion_matrix(labels, predictions).numpy()
        # Get number of all true positives in each class
        all_true_positives = np.diag(confusion_matrix)
        # Get number of all elements in each class
        all_class_sum = np.sum(confusion_matrix, axis=1)
        # Get rid of classes that don't show in batch
        zero_index = np.where(all_class_sum == 0)[0]
        all_true_positives = np.delete(all_true_positives, zero_index)
        all_class_sum = np.delete(all_class_sum, zero_index)
    
        print("confusion_matrix:\n {},\n all_true_positives:\n {},\n all_class_sum:\n {}".format(
                                                confusion_matrix, all_true_positives, all_class_sum))
        # Average TruePositives / TotalElements wrt all classes that show in batch
        return np.mean(all_true_positives / all_class_sum)
    
    avg_recall = average_recall(y_t, y_p)
    print(avg_recall)
    

    输出:

    confusion_matrix:
     [[1 0 0 0]
     [1 1 0 0]
     [0 0 0 0]
     [0 0 0 2]],
     all_true_positives:
     [1 1 2],
     all_class_sum:
     [1 2 2]
    0.8333333333333334
    

    仅使用 tensorflow 实现:

    import tensorflow as tf
    
    y_t = tf.constant([[1, 0, 0, 0], [0, 1, 0, 0], [0, 1, 0, 0], [0, 0, 0, 1], [0, 0, 0, 1]], dtype=tf.float32)
    y_p = tf.constant([[1, 0, 0, 0], [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 0, 1], [0, 0, 0, 1]], dtype=tf.float32)
    
    def average_recall(y_true, y_pred):
        # Get indexes of both labels and predictions
        labels = tf.argmax(y_true, axis=1)
        predictions = tf.argmax(y_pred, axis=1)
        # Get confusion matrix from labels and predictions
        confusion_matrix = tf.math.confusion_matrix(labels, predictions)
        # Get number of all true positives in each class
        all_true_positives = tf.linalg.diag_part(confusion_matrix)
        # Get number of all elements in each class
        all_class_sum = tf.reduce_sum(confusion_matrix, axis=1)
        # Get rid of classes that don't show in batch
        mask = tf.not_equal(all_class_sum, tf.constant(0))
        all_true_positives = tf.boolean_mask(all_true_positives, mask)
        all_class_sum = tf.boolean_mask(all_class_sum, mask)
    
        print("confusion_matrix:\n {},\n all_true_positives:\n {},\n all_class_sum:\n {}".format(
                                                confusion_matrix, all_true_positives, all_class_sum))
        # Average TruePositives / TotalElements wrt all classes that show in batch
        return tf.reduce_mean(all_true_positives / all_class_sum)
    
    avg_recall = average_recall(y_t, y_p)
    print(avg_recall)
    

    输出:

    confusion_matrix:
     [[1 0 0 0]
     [1 1 0 0]
     [0 0 0 0]
     [0 0 0 2]],
     all_true_positives:
     [1 1 2],
     all_class_sum:
     [1 2 2]
    tf.Tensor(0.8333333333333334, shape=(), dtype=float64)
    

    参考:

    tf.math.confusion_matrix

    Calculate precision and recall for multiclass classification using confusion matrix

    【讨论】:

    • 是的!那很完美。我的问题不是 numpy 实现,我的问题是 keras 实现。我找不到允许我这样做的 keras.backend 函数。我会回答我的问题
    • @LuisFelipe 似乎您对答案很满意,但请明确一点,我的答案是否解决了您的问题,或者您想要 keras 实现?
    • 我想要一个 keras 实现。我将此指标与急切模式一起使用并将张量转换为 numpy 数组,但这比较慢。 Keras 的实现会很棒
    • @我也添加了tf 实现,查看答案!
    • 我所做的一切都是平等的,但我没有找到任何与 np.delete 等效的东西。你做得很好,谢谢!@Mr.ForExample
    猜你喜欢
    • 1970-01-01
    • 2020-04-12
    • 2020-10-08
    • 1970-01-01
    • 2021-06-03
    • 2020-02-05
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多