【问题标题】:Base64 images with Keras and Google Cloud ML使用 Keras 和 Google Cloud ML 的 Base64 图像
【发布时间】:2018-06-08 09:56:00
【问题描述】:

我正在使用 Keras 预测图像类。它适用于 Google Cloud ML (GCML),但为了提高效率,需要将其更改为传递 base64 字符串而不是 json 数组。 Related Documentation

我可以轻松运行 python 代码将 base64 字符串解码为 json 数组,但是在使用 GCML 时,我没有机会运行预处理步骤(除非可能在 Keras 中使用 Lambda 层,但我不认为这是正确的方法)。

Another answer 建议添加类型为tf.string 的tf.placeholder,这是有道理的,但是如何将其合并到 Keras 模型中呢?

这里是训练模型和为 GCML 保存导出模型的完整代码...

import os
import numpy as np
import tensorflow as tf
import keras
from keras import backend as K
from keras.models import Sequential
from keras.layers import Dense, Dropout, Flatten
from keras.layers import Conv2D, MaxPooling2D
from keras.preprocessing import image
from tensorflow.python.platform import gfile

IMAGE_HEIGHT = 138
IMAGE_WIDTH = 106
NUM_CLASSES = 329

def preprocess(filename):
    # decode the image file starting from the filename
    # end up with pixel values that are in the -1, 1 range
    image_contents = tf.read_file(filename)
    image = tf.image.decode_png(image_contents, channels=1)
    image = tf.image.convert_image_dtype(image, dtype=tf.float32) # 0-1
    image = tf.expand_dims(image, 0) # resize_bilinear needs batches
    image = tf.image.resize_bilinear(image, [IMAGE_HEIGHT, IMAGE_WIDTH], align_corners=False)
    image = tf.subtract(image, 0.5)
    image = tf.multiply(image, 2.0) # -1 to 1
    image = tf.squeeze(image,[0])
    return image



filelist = gfile.ListDirectory("images")
sess = tf.Session()
with sess.as_default():
    x = np.array([np.array(     preprocess(os.path.join("images", filename)).eval()      ) for filename in filelist])

input_shape = (IMAGE_HEIGHT, IMAGE_WIDTH, 1)   # 1, because preprocessing made grayscale

# in our case the labels come from part of the filename
y = np.array([int(filename[filename.index('_')+1:-4]) for filename in filelist])
# convert class labels to numbers
y = keras.utils.to_categorical(y, NUM_CLASSES)

########## TODO: something here? ##########
image = K.placeholder(shape=(), dtype=tf.string)
decoded = tf.image.decode_jpeg(image, channels=3)
# scores = build_model(decoded)


model = Sequential()

# model.add(decoded)

model.add(Conv2D(32, kernel_size=(2, 2), activation='relu', input_shape=input_shape))
model.add(Conv2D(64, (3, 3), activation='relu'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model.add(Dropout(0.25))
model.add(Flatten())
model.add(Dense(64, activation='relu'))
model.add(Dropout(0.25))
model.add(Dense(num_classes, activation='softmax'))

model.compile(loss=keras.losses.categorical_crossentropy,
            optimizer=keras.optimizers.Adadelta(),
            metrics=['accuracy'])

model.fit(
    x,
    y,
    batch_size=64,
    epochs=20,
    verbose=1,
    validation_split=0.2,
    shuffle=False
    )

predict_signature = tf.saved_model.signature_def_utils.build_signature_def(
    inputs={'input_bytes':tf.saved_model.utils.build_tensor_info(model.input)},
    ########## TODO: something here? ##########
    # inputs={'input': image },    # input name must have "_bytes" suffix to use base64.
    outputs={'formId': tf.saved_model.utils.build_tensor_info(model.output)},
    method_name=tf.saved_model.signature_constants.PREDICT_METHOD_NAME
)

builder = tf.saved_model.builder.SavedModelBuilder("exported_model")

builder.add_meta_graph_and_variables(
    sess=K.get_session(),
    tags=[tf.saved_model.tag_constants.SERVING],
    signature_def_map={
        tf.saved_model.signature_constants.DEFAULT_SERVING_SIGNATURE_DEF_KEY: predict_signature
    },
    legacy_init_op=tf.group(tf.tables_initializer(), name='legacy_init_op')
)

builder.save()

这和我的previous question有关。

更新:

问题的核心是如何将调用 decode 的占位符合并到 Keras 模型中。换句话说,在创建将 base64 字符串解码为张量的占位符之后,如何将其合并到 Keras 运行的内容中?我认为它需要是一个层。

image = K.placeholder(shape=(), dtype=tf.string)
decoded = tf.image.decode_jpeg(image, channels=3)
model = Sequential()

# Something like this, but this fails because it is a tensor, not a Keras layer.  Possibly this is where a Lambda layer comes in?
model.add(decoded)
model.add(Conv2D(32, kernel_size=(2, 2), activation='relu', input_shape=input_shape))
...

更新 2:

正在尝试使用 lambda 层来完成此操作...

import keras
from keras.models import Sequential
from keras.layers import Lambda
from keras import backend as K
import tensorflow as tf

image = K.placeholder(shape=(), dtype=tf.string)
model = Sequential()
model.add(Lambda(lambda image: tf.image.decode_jpeg(image, channels=3), input_shape=() ))

给出错误:TypeError: Input 'contents' of 'DecodeJpeg' Op has type float32 that does not match expected type of string.

【问题讨论】:

  • 根据您的第二次更新,显示类型之间不匹配。检查您的 decode 函数是否返回您期望的数据类型,或者尝试将占位符 dtype 更改为 float。
  • 你找到解决办法了吗?
  • 我添加了 dtype 但是我得到了这个错误 Shape must be rank 0 but is rank 2 for 'lambda_4/DecodePng' (op: 'DecodePng') with input shapes: [?,?].

标签: keras google-cloud-ml


【解决方案1】:

首先我使用 tf.keras 但这应该不是什么大问题。 下面是一个如何读取 base64 解码的 jpeg 的示例:

def preprocess_and_decode(img_str, new_shape=[299,299]):
    img = tf.io.decode_base64(img_str)
    img = tf.image.decode_jpeg(img, channels=3)
    img = tf.image.resize_images(img, new_shape, method=tf.image.ResizeMethod.BILINEAR, align_corners=False)
    # if you need to squeeze your input range to [0,1] or [-1,1] do it here
    return img
InputLayer = Input(shape = (1,),dtype="string")
OutputLayer = Lambda(lambda img : tf.map_fn(lambda im : preprocess_and_decode(im[0]), img, dtype="float32"))(InputLayer)
base64_model = tf.keras.Model(InputLayer,OutputLayer)   

上面的代码创建了一个模型,该模型采用任意大小的 jpeg,将其调整为 299x299 并返回为 299x299x3 张量。此模型可以直接导出到 saved_model 并用于 Cloud ML Engine 服务。这有点愚蠢,因为它唯一做的就是将 base64 转换为张量。

如果您需要将此模型的输出重定向到现有训练和编译模型(例如 inception_v3)的输入,您必须执行以下操作:

base64_input = base64_model.input
final_output = inception_v3(base64_model.output)
new_model = tf.keras.Model(base64_input,final_output)

这个 new_model 可以保存。它采用 base64 jpeg 并返回由 inception_v3 部分标识的类。

【讨论】:

    【解决方案2】:

    另一个答案建议添加tf.placeholder 类型为tf.string,这是有道理的,但是如何将其合并到 Keras 模型中?

    在 Keras 中,您可以通过以下方式访问您选择的后端(在本例中为 Tensorflow):

    from keras import backend as K
    

    你似乎已经在你的代码中导入了这个。这将使您能够访问您选择的后端可用的一些本机方法和资源。 Keras 后端包含一种用于创建占位符的方法,以及其他实用程序。关于占位符,我们可以看到 Keras docs 对它们的说明:

    占位符

    keras.backend.placeholder(shape=None, ndim=None, dtype=None, sparse=False, name=None)

    实例化一个占位符张量并返回它。

    它还给出了一些使用示例:

    >>> from keras import backend as K
    >>> input_ph = K.placeholder(shape=(2, 4, 5))
    >>> input_ph._keras_shape
    (2, 4, 5)
    >>> input_ph
    <tf.Tensor 'Placeholder_4:0' shape=(2, 4, 5) dtype=float32>
    

    如您所见,这将返回一个 Tensorflow 张量,其形状为 (2,4,5),dtype 为浮点数。如果你在做这个例子时有另一个后端,你会得到另一个张量对象(肯定是 Theano 对象)。因此,您可以使用此 placeholder() 来调整您之前的 question 上的 solution。

    总之,您可以使用导入为K(或任何您想要的)的后端来调用您选择的后端可用的方法和对象,方法是在所需方法上执行K.foo.bar()。我建议您阅读 Keras Backend 的内容,以探索更多在未来情况下对您有用的东西。

    更新:根据您的编辑。是的,这个占位符应该是模型中的一个层。具体来说,它应该是模型的输入层,因为它保存解码后的图像(因为 Keras 需要它)以进行分类。

    【讨论】:

    • 感谢您的回复。我了解 K.placeholder 提供了创建 tf.placeholder 的访问权限。然而,调整代码(简单地将 tf.placeholder 替换为 K.placeholder)存在问题。首先,tf.image.decode_jpeg 期望图像的等级为 0,但它被定义为等级 1(因为 shape=[None])。回答者如何使用 'scores' 变量来保存模型构建的结果以及模型签名的输出也让他们感到困惑。你能包括具体的代码吗?
    • @user3567174 肯定会在我明天早上使用我的电脑时完成,因为我目前正在移动,无法浏览和尝试太多
    • @user3567174 你试过我的建议了吗?我想知道如果那里有您不理解的东西,为什么您会接受该答案-“ ...期望图像为 0 级,但它被定义为 1 级(因为 shape=[None])。” - 在这种情况下,不要传递[None](因为你不想要批次)而是传递相应的形状。您是否尝试过 shape=() 而不是表示排名 0?...我不太明白该用户在那里对您的回答,但似乎 scores 存储了从您的模型的其余部分获得的分类(我理解回答者总结为build_model)。
    • 我没有意识到你可以创建一个等级为 0 的张量;似乎是矛盾的。在尝试实施之前,我接受了您对另一个问题的回答,因为它似乎解决了我的问题。我已经用更多细节更新了这个问题。如果这不能说明我在问什么,请告诉我。
    • @user3567174 根据您的更新,是的,它应该是一个层。它是您的输入层,因为它携带您的解码图像进行分类。更新答案以反映这一点。作为旁注,我通常更喜欢不使用顺序模型并选择Functional API,因为我发现它可以让您进行更多自定义,所以也许您也可以查看一下你更喜欢它:)
    猜你喜欢
    • 2018-05-03
    • 1970-01-01
    • 2019-10-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2018-09-19
    • 2017-08-03
    • 1970-01-01
    相关资源
    最近更新 更多