【问题标题】:How to store a dictionary and map words to ints when using Tensorflow Serving?使用 Tensorflow Serving 时如何存储字典并将单词映射到整数?
【发布时间】:2017-11-20 18:58:44
【问题描述】:

我已经在 Tensorflow 上训练了一个 LSTM RNN 分类模型。我正在保存和恢复检查点以重新训练和使用模型进行测试。现在我想使用 Tensorflow 服务,以便在生产中使用该模型。

最初,我会解析一个语料库以创建我的字典,然后将其用于将字符串中的单词映射到整数。然后我会将这个字典存储在一个 pickle 文件中,当恢复检查点和重新训练数据集或仅用于使用模型时可以重新加载该文件,以便映射是一致的。使用 SavedModelBuilder 保存模型时如何存储此字典?

我的神经网络代码如下。保存模型的代码接近尾声(我包括上下文的整个结构的概述):

...


# Read files and store them in variables
with open('./someReview.txt', 'r') as f:
    reviews = f.read()
with open('./someLabels.txt', 'r') as f:
    labels = f.read()

...

#Pre-processing functions
#Parse through dataset and create a vocabulary
vocab_to_int, reviews = RnnPreprocessing.map_vocab_to_int(reviews)
with open(pickle_path, 'wb') as handle:
    pickle.dump(vocab_to_int, handle, protocol=pickle.HIGHEST_PROTOCOL)

#More preprocessing functions
...


# Building the graph
lstm_size = 256
lstm_layers = 2
batch_size = 1000
learning_rate = 0.01            
n_words = len(vocab_to_int) + 1 

# Create the graph object
tf.reset_default_graph()
with tf.name_scope('inputs'):
    inputs_ = tf.placeholder(tf.int32, [None, None], name="inputs")
    labels_ = tf.placeholder(tf.int32, [None, None], name="labels")
    keep_prob = tf.placeholder(tf.float32, name="keep_prob")

#Create embedding layer LSTM cell, LSTM Layers

...

# Forward pass
with tf.name_scope("RNN_forward"):
    outputs, final_state = tf.nn.dynamic_rnn(cell, embed, initial_state=initial_state)


# Output. We are only interested in the latest output of the lstm cell
with tf.name_scope('predictions'):
    predictions = tf.contrib.layers.fully_connected(outputs[:, -1], 1, activation_fn=tf.sigmoid)
    tf.summary.histogram('predictions', predictions)
#More functions for cost, accuracy, optimizer initialization

... 

# Training
epochs = 1
with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    iteration = 1
    for e in range(epochs):
        state = sess.run(initial_state)

        for ii, (x, y) in enumerate(get_batches(train_x, train_y, batch_size), 1):
            feed = {inputs_: x,
                    labels_: y[:, None],
                    keep_prob: 0.5,
                    initial_state: state}
            summary, loss, state, _ = sess.run([merged, cost, final_state, optimizer], feed_dict=feed)

            train_writer.add_summary(summary, iteration)

            if iteration%1==0:
                print("Epoch: {}/{}".format(e, epochs),
                      "Iteration: {}".format(iteration),
                      "Train loss: {:.3f}".format(loss))

            if iteration%2==0:
                val_acc = []
                val_state = sess.run(cell.zero_state(batch_size, tf.float32))
                for x, y in get_batches(val_x, val_y, batch_size):
                    feed = {inputs_: x,
                            labels_: y[:, None],
                            keep_prob: 1,
                            initial_state: val_state}
                    summary, batch_acc, val_state = sess.run([merged, accuracy, final_state], feed_dict=feed)
                    val_acc.append(batch_acc)
                print("Val acc: {:.3f}".format(np.mean(val_acc)))
            iteration +=1
            test_writer.add_summary(summary, iteration)



    #Saving the model
    export_path = './SavedModel'
    print ('Exporting trained model to %s'%(export_path))

    builder = saved_model_builder.SavedModelBuilder(export_path)

    # Build the signature_def_map.    
    classification_inputs = utils.build_tensor_info(inputs_)
    classification_outputs_classes = utils.build_tensor_info(labels_)

    classification_signature = signature_def_utils.build_signature_def(
        inputs={signature_constants.CLASSIFY_INPUTS: classification_inputs},
        outputs={
          signature_constants.CLASSIFY_OUTPUT_CLASSES:
              classification_outputs_classes,
        },
      method_name=signature_constants.CLASSIFY_METHOD_NAME)


    legacy_init_op = tf.group(
        tf.tables_initializer(), name='legacy_init_op')
    #add the sigs to the servable
    builder.add_meta_graph_and_variables(
        sess, [tag_constants.SERVING],
        signature_def_map={
            signature_constants.DEFAULT_SERVING_SIGNATURE_DEF_KEY:
                classification_signature
        },
        legacy_init_op=legacy_init_op)
    print ("added meta graph and variables")

    #save it!
    builder.save()
    print("model saved")

我不完全确定这是否是保存此类模型的正确方法,但这是我在文档和在线教程中找到的唯一实现。

我没有找到任何示例或任何明确的指南来保存字典或在文档中恢复已保存模型时如何使用它。

使用检查点时,我会在运行会话之前加载 pickle 文件。如何恢复此 savedModel 以便我可以使用字典使用相同的单词到 int 映射?有什么具体的方法可以保存或加载模型吗?

我还添加了 inputs_ 作为输入签名的输入。这是单词被映射后的整数序列。我无法将字符串指定为输入,因为我得到了 AttributeError: 'str' object has no attribute 'dtype' 。在这种情况下,单词在生产中的模型中究竟是如何映射到整数的?

【问题讨论】:

    标签: machine-learning tensorflow lstm tensorflow-serving word-embedding


    【解决方案1】:

    使用tf.feature_column 中的实用程序实现您的预​​处理,并且在服务中使用相同的整数映射将很简单。

    【讨论】:

    • 即便如此,当您将输入中的所有单词映射到整数时,您仍然需要一个可以参考的词汇或词汇表,除非我遗漏了什么。你会在哪里存储这本词典。能详细点吗?
    【解决方案2】:

    解决此问题的一种方法是将词汇表存储在模型的图表中。这将与模型一起提供。

    ...
    
    
    vocab_table = lookup.index_table_from_file(vocabulary_file='data/vocab.csv', num_oov_buckets=1, default_value=-1)
    text = features[commons.FEATURE_COL]
    words = tf.string_split(text)
    dense_words = tf.sparse_tensor_to_dense(words, default_value=commons.PAD_WORD)
    word_ids = vocab_table.lookup(dense_words)
    
    padding = tf.constant([[0, 0], [0, commons.MAX_DOCUMENT_LENGTH]])
    # Pad all the word_ids entries to the maximum document length
    word_ids_padded = tf.pad(word_ids, padding)
    word_id_vector = tf.slice(word_ids_padded, [0, 0], [-1, commons.MAX_DOCUMENT_LENGTH])
    

    来源:https://github.com/KishoreKarunakaran/CloudML-Serving/blob/master/text/imdb_cnn/model/cnn_model.py#L83

    【讨论】:

      猜你喜欢
      • 2019-03-18
      • 1970-01-01
      • 2020-07-14
      • 1970-01-01
      • 2019-07-27
      • 2018-09-11
      • 1970-01-01
      • 1970-01-01
      • 2015-10-17
      相关资源
      最近更新 更多