【问题标题】:Triplet model for image retrieval from the Keras pretrained network从 Keras 预训练网络中检索图像的三元组模型
【发布时间】:2018-04-18 03:55:00
【问题描述】:

我想实现一个图像检索模型。该模型将使用三元组损失函数(与 facenet 或类似架构相同)进行训练。我的想法是使用来自 Keras 的预训练分类模型(例如 resnet50),并使其成为三重架构。这是我在 Keras 中的模型:

resnet_input = Input(shape=(224,224,3))
resnet_model = ResNet50(weights='imagenet', include_top = False, input_tensor=resnet_input)
net = resnet_model.output

net = Flatten(name='flatten')(net) 
net = Dense(512, activation='relu', name='embded')(net)
net = Lambda(l2Norm, output_shape=[512])(net)

base_model = Model(resnet_model.input, net, name='resnet_model')

input_shape=(224,224,3)
input_anchor = Input(shape=input_shape, name='input_anchor')
input_positive = Input(shape=input_shape, name='input_pos')
input_negative = Input(shape=input_shape, name='input_neg')

net_anchor = base_model(input_anchor)
net_positive = base_model(input_positive)
net_negative = base_model(input_negative)

positive_dist = Lambda(euclidean_distance, name='pos_dist')([net_anchor, net_positive])
negative_dist = Lambda(euclidean_distance, name='neg_dist')([net_anchor, net_negative])

stacked_dists = Lambda( 
            lambda vects: K.stack(vects, axis=1),
            name='stacked_dists'
)([positive_dist, negative_dist])

model = Model([input_anchor, input_positive, input_negative], stacked_dists, name='triple_siamese')

def triplet_loss(_, y_pred):
    margin = K.constant(1)
    return K.mean(K.maximum(K.constant(0), K.square(y_pred[0]) - K.square(y_pred[1]) + margin))

def accuracy(_, y_pred):
    return K.mean(y_pred[0] < y_pred[1])

def l2Norm(x):
    return  K.l2_normalize(x, axis=-1)

def euclidean_distance(vects):
    x, y = vects
    return K.sqrt(K.maximum(K.sum(K.square(x - y), axis=1, keepdims=True), K.epsilon()))

模型应该为每张图片预测一个特征向量。如果图像来自同一类,则这些向量之间的距离(在本例中为欧几里得)应该接近于零,如果不是,则接近于一。

我已经尝试过不同的学习步骤、批量大小、损失函数中的不同边距、从原始 resnet 模型中选择不同的输出层、在 resnet 的末尾添加不同的层、只训练新添加的层 vs 训练整个层模型。我也尝试使用这个没有预训练权重的 resnet 模型,无论我做什么,结果还是一样,大约 0.5 的准确率和 1.0 的损失。 (输入图像以 keras.applications.resnet50.preprocess_input 对该模型的预期方式进行预处理)

我没有进行任何可能导致收敛缓慢的硬负挖掘,但在这种情况下,0.5 的准确度(检查函数)仍然是随机预测。

所以我开始想,也许我错过了一些非常重要的东西(这是一个非常困难的架构)。因此,如果您在我的实施中发现有问题或可疑之处,我会非常高兴。

【问题讨论】:

    标签: python tensorflow keras conv-neural-network


    【解决方案1】:

    如果有人感兴趣,重写

    y_pred[0] 和y_pred[1]

    到

    y_pred[:,0,0] 和 y_pred[:,1,0]

    修复它。

    现在模型似乎在训练(损失在减少,准确率在增加)。

    【讨论】:

    • 三联网络是否有“准确性”?另外,如果你解释一下为什么 y_pred[0] 变成了 y_pred[:,0,0],我将不胜感激。
    • @LKM 好吧,这不是通常的准确性。这个告诉我有多少图像三元组被正确分类(来自同一类的两张图像,一张来自不同类的图像)。它至少在说明一些事情。为什么我做了我不确定的改变。我尝试了很多东西,这一个奏效了......
    【解决方案2】:

    我没有足够的声望点来发表评论,所以我这样写。

    我想通过使用您的代码来做与您正在做的事情类似的事情。

    我是 CNN 的新手,我不确定我的训练数据应该是什么样子。你愿意分享你剩下的代码吗?我将不胜感激!

    编辑:

    为了回答我自己可能对某人有用的问题,这就是我在假期照片 (http://lear.inrialpes.fr/%7Ejegou/data.php) 上所做的,它有效:

    def get_random_image(img_groups, group_names, gid):
        gname = group_names[gid]
        photos = img_groups[gname]
        pid = np.random.choice(np.arange(len(photos)), size=1)[0]
        pname = photos[pid]
        return gname + pname + ".jpg"
    
    def create_triples(image_dir):
        img_groups = {}
        for img_file in os.listdir(image_dir):
            prefix, suffix = img_file.split(".")
            gid, pid = prefix[0:4], prefix[4:]
    
            if gid in img_groups.keys():
                img_groups[gid].append(pid)
            else:
                img_groups[gid] = [pid]
        pos_triples, neg_triples = [], []
    
        for key in img_groups.keys():
            triples = [(key + x[0] + ".jpg", key + x[1] + ".jpg", str(int(key)+3 if int(key)<1495 else int(key)-3)+'01'+'.jpg')
                     for x in itertools.combinations(img_groups[key], 2)]
            pos_triples.extend(triples)
    
        return pos_triples
    
    def triplet_loss(y_true, y_pred):
            margin = K.constant(0.2)
            return K.mean(K.maximum(K.constant(0), K.square(y_pred[:,0,0]) - K.square(y_pred[:,1,0]) + margin))
    
    def accuracy(y_true, y_pred):
        return K.mean(y_pred[:,0,0] < y_pred[:,1,0])
    
    def l2Norm(x):
        return  K.l2_normalize(x, axis=-1)
    
    def euclidean_distance(vects):
        x, y = vects
        return K.sqrt(K.maximum(K.sum(K.square(x - y), axis=1, keepdims=True), K.epsilon()))
    
    triples_data = create_triples(IMAGE_DIR)
    
    
    dim = 1500
    h = 299
    w= 299
    anchor =np.zeros((dim,h,w,3))
    positive =np.zeros((dim,h,w,3))
    negative =np.zeros((dim,h,w,3))
    
    
    for n,val in enumerate(triples_data[0:1500]):
        image_anchor = plt.imread(os.path.join(IMAGE_DIR, val[0]))
        image_anchor = imresize(image_anchor, (h, w))    
        image_anchor = image_anchor.astype("float32")
        #image_anchor = image_anchor/255.
        image_anchor = keras.applications.resnet50.preprocess_input(image_anchor, data_format='channels_last')
        anchor[n] = image_anchor
    
        image_positive = plt.imread(os.path.join(IMAGE_DIR, val[1]))
        image_positive = imresize(image_positive, (h, w))
        image_positive = image_positive.astype("float32")
        #image_positive = image_positive/255.
        image_positive = keras.applications.resnet50.preprocess_input(image_positive, data_format='channels_last')
        positive[n] = image_positive
    
        image_negative = plt.imread(os.path.join(IMAGE_DIR, val[2]))
        image_negative = imresize(image_negative, (h, w))
        image_negative = image_negative.astype("float32")
        #image_negative = image_negative/255.
        image_negative = keras.applications.resnet50.preprocess_input(image_negative, data_format='channels_last')
        negative[n] = image_negative
    
    Y_train = np.random.randint(2, size=(1,2,dim)).T
    
    
    resnet_input = Input(shape=(h,w,3))
    resnet_model = ResNet50(weights='imagenet', include_top = False, input_tensor=resnet_input)
    
    
    for layer in resnet_model.layers:
        layer.trainable = False  
    
    
    net = resnet_model.output
    net = Flatten(name='flatten')(net) 
    net = Dense(128, activation='relu', name='embed')(net)
    net = Dense(128, activation='relu', name='embed2')(net)
    net = Dense(128, activation='relu', name='embed3')(net)
    net = Lambda(l2Norm, output_shape=[128])(net)
    
    base_model = Model(resnet_model.input, net, name='resnet_model')
    
    input_shape=(h,w,3)
    input_anchor = Input(shape=input_shape, name='input_anchor')
    input_positive = Input(shape=input_shape, name='input_pos')
    input_negative = Input(shape=input_shape, name='input_neg')
    
    net_anchor = base_model(input_anchor)
    net_positive = base_model(input_positive)
    net_negative = base_model(input_negative)
    
    positive_dist = Lambda(euclidean_distance, name='pos_dist')([net_anchor, net_positive])
    negative_dist = Lambda(euclidean_distance, name='neg_dist')([net_anchor, net_negative])
    
    stacked_dists = Lambda( 
                lambda vects: K.stack(vects, axis=1),
                name='stacked_dists'
    )([positive_dist, negative_dist])
    
    
    model = Model([input_anchor, input_positive, input_negative], stacked_dists, name='triple_siamese')
    
    model.compile(optimizer="rmsprop", loss=triplet_loss, metrics=[accuracy])
    
    model.fit([anchor, positive, negative], Y_train, epochs=50,  batch_size=15, validation_split=0.2)
    
    model.save('triplet_loss_resnet50.h5')
    

    【讨论】:

    • 为什么 "K.square(y_pred[0]) - K.square(y_pred[1]) + margin" 变成 "K.square(y_pred[:,0,0]) - K.square(y_pred[:,1,0]) + 边距"?我不明白。如果您能解释您的代码,我将不胜感激。
    • 使用上面的代码,我有时会因为丢失而得到 0。有什么我想念的吗?
    • @第一个问题:因为模型的输出形状是(?, 2, 1)。您想减去整个训练集的平方作为锚点和正数之间的距离(即 [:, 0, 0]),以及整个训练集的平方作为锚点和负数之间的距离(即 [:, 1, 0 ])。
    • @第二个问题:没有更多信息,我很难说什么。对我来说,这段代码有效。也许训练对没有正确生成,所以你实际上是在比较之间的距离,例如锚+正和锚+正?
    • model.fit 中的 Y_train 是什么?它的形状是什么?如果已经指定loss=triplet loss,那么需要什么lamba层。请解释
    猜你喜欢
    • 2021-03-14
    • 2019-06-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2020-04-16
    • 1970-01-01
    • 1970-01-01
    • 2018-04-15
    相关资源
    最近更新 更多