【问题标题】:Keras Text Preprocessing - Saving Tokenizer object to file for scoringKeras 文本预处理 - 将 Tokenizer 对象保存到文件以进行评分
【发布时间】:2017-11-29 05:11:13
【问题描述】:

我已经按照以下步骤(大致)使用 Keras 库训练了一个情绪分类器模型。

  1. 使用 Tokenizer 对象/类将文本语料库转换为序列
  2. 使用 model.fit() 方法构建模型
  3. 评估此模型

现在,为了使用此模型进行评分,我能够将模型保存到文件并从文件中加载。但是我还没有找到将 Tokenizer 对象保存到文件的方法。如果没有这个,我每次需要对一个句子进行评分时都必须处理语料库。有没有办法解决这个问题?

【问题讨论】:

    标签: machine-learning neural-network nlp deep-learning keras


    【解决方案1】:

    最常用的方法是使用pickle 或joblib。这里有一个关于如何使用pickle 来保存Tokenizer 的示例:

    import pickle
    
    # saving
    with open('tokenizer.pickle', 'wb') as handle:
        pickle.dump(tokenizer, handle, protocol=pickle.HIGHEST_PROTOCOL)
    
    # loading
    with open('tokenizer.pickle', 'rb') as handle:
        tokenizer = pickle.load(handle)
    

    【讨论】:

    • 你会在测试集上再次调用 tokenizer.fit_on_texts 吗?
    • 没有。如果您再次调用 fit* 它可能会更改索引。已加载 pickle 的标记器已准备好使用。
    • 等等。您必须同时保存模型和标记器以便将来运行模型?
    • 当然!它们有 2 个不同的角色,tokenizer 会将文本转换为向量,重要的是在训练和测试之间拥有相同的向量空间。
    【解决方案2】:

    Tokenizer 类具有将日期保存为 JSON 格式的功能:

    tokenizer_json = tokenizer.to_json()
    with io.open('tokenizer.json', 'w', encoding='utf-8') as f:
        f.write(json.dumps(tokenizer_json, ensure_ascii=False))
    

    可以使用keras_preprocessing.text中的tokenizer_from_json函数加载数据:

    with open('tokenizer.json') as f:
        data = json.load(f)
        tokenizer = tokenizer_from_json(data)
    

    【讨论】:

    • tokenizer_from_json 似乎不再在 Keras 中可用,或者更确切地说,它没有在他们的文档中列出或在 conda @Max 的包中可用,你仍然这样做吗?
    • @benbyford 我使用来自 PyPI 的 Keras-Preprocessing==1.0.9 包和函数 is avaiable
    • tokenizer_to_json 应该很快就会在 tensorflow > 2.0.0 上可用,看到这个pr 同时from keras_preprocessing.text import tokenizer_from_json 可以使用
    • 这对我有用。谢谢
    【解决方案3】:

    接受的答案清楚地展示了如何保存分词器。以下是对(一般)在拟合或保存后后评分问题的评论。假设列表texts 由两个列表Train_text 和Test_text 组成,其中Test_text 中的标记集是Train_text 中标记集的子集(乐观假设)。然后fit_on_texts(Train_text) 为texts_to_sequences(Test_text) 提供不同的结果,与先调用fit_on_texts(texts) 然后text_to_sequences(Test_text) 相比。

    具体例子:

    from keras.preprocessing.text import Tokenizer
    
    docs = ["A heart that",
             "full up like",
             "a landfill",
            "no surprises",
            "and no alarms"
             "a job that slowly"
             "Bruises that",
             "You look so",
             "tired happy",
             "no alarms",
            "and no surprises"]
    docs_train = docs[:7]
    docs_test = docs[7:]
    # EXPERIMENT 1: FIT  TOKENIZER ONLY ON TRAIN
    T_1 = Tokenizer()
    T_1.fit_on_texts(docs_train)  # only train set
    encoded_train_1 = T_1.texts_to_sequences(docs_train)
    encoded_test_1 = T_1.texts_to_sequences(docs_test)
    print("result for test 1:\n%s" %(encoded_test_1,))
    
    # EXPERIMENT 2: FIT TOKENIZER ON BOTH TRAIN + TEST
    T_2 = Tokenizer()
    T_2.fit_on_texts(docs)  # both train and test set
    encoded_train_2 = T_2.texts_to_sequences(docs_train)
    encoded_test_2 = T_2.texts_to_sequences(docs_test)
    print("result for test 2:\n%s" %(encoded_test_2,))
    

    结果:

    result for test 1:
    [[3], [10, 3, 9]]
    result for test 2:
    [[1, 19], [5, 1, 4]]
    

    当然,如果不满足上述乐观假设,并且Test_text中的token集合与Train_test的token集合不相交,那么test 1会产生一个空括号列表[].

    【讨论】:

    • 故事的寓意:如果使用词嵌入和 keras 的 Tokenizer,在非常大的语料库上只使用一次 fit_on_texts;或改用字符 n-gram。
    • 我不明白您要传达的信息是什么:为什么一开始就适合测试文档?根据定义,无论您在做什么,测试都必须保存在保险库中,就好像您一开始就不知道自己拥有它一样。
    • @gented:您可能会将无监督文本解析与有监督机器学习混淆。如果我错了,请纠正我,但 keras 的 Tokenizer 没有附加用于泛化的损失函数;因此,这不是(监督)机器学习问题——这似乎是您的假设。我试图传达的信息总结在我上面的第一条评论中(“故事的道德......”),可能值得重读。
    • @gented 优点。抱歉,如果命名法让您感到困惑;我在接受的答案中与 cmets 保持一定的一致性。
    • 我同意@gented 的观点,因为您不想将标记器放入测试集中,因为这样您在测试时就消除了 oov 标记的可能性,从而违背了测试集的目的。这不是因为分词器有损失,而是因为测试集中的数据泄漏到您的训练数据中。
    【解决方案4】:

    我在 keras Repo 中创建了问题 https://github.com/keras-team/keras/issues/9289。在更改 API 之前,该问题有一个指向 gist 的链接,该链接具有代码来演示如何保存和恢复标记器,而无需标记器适合的原始文档。我更喜欢将我的所有模型信息存储在 JSON 文件中(因为原因,但主要是混合 JS/Python 环境),这将允许这样做,即使使用 sort_keys=True

    【讨论】:

    • 链接的要点看起来像是“重新加载”训练有素的标记器的好方法。但是,最初的问题可能与将先前保存的分词器“扩展”到新的(测试)文本有关;这部分似乎仍然是开放的(否则,如果模型不用于“评分”新数据,为什么还要“保存”它?)
    • 我认为他们的意图很明确“没有这个,我每次需要对一个句子进行评分时都必须处理语料库”。据此,我推测他们希望跳过标记化步骤并在其他数据上评估经过训练的模型。他们不问任何其他问题,那就是您所期待的。他们和大多数人一样,只想在大多数教程中跳过的不同数据集上使用先前标记化的数据。因此,我认为我的回答 1)回答了所问的问题,并且 2)提供了工作代码。
    • 公平点。问题是“将 Tokenizer 对象保存到文件以进行评分”,因此人们可能会认为他们也在询问评分(可能是新数据)。
    【解决方案5】:

    我在link by @thusv89 处找到了以下 sn-p。

    保存对象:

    import pickle
    
    with open('data_objects.pickle', 'wb') as handle:
        pickle.dump(
            {'input_tensor': input_tensor, 
             'target_tensor': target_tensor, 
             'inp_lang': inp_lang,
             'targ_lang': targ_lang,
            }, handle, protocol=pickle.HIGHEST_PROTOCOL)
    

    加载对象:

    with open("dataset_fr_en.pickle", 'rb') as f:
        data = pickle.load(f)
        input_tensor = data['input_tensor']
        target_tensor = data['target_tensor']
        inp_lang = data['inp_lang']
        targ_lang = data['targ_lang']
    

    【讨论】:

      【解决方案6】:

      很简单,因为 Tokenizer 类提供了保存和加载两个函数:

      保存——Tokenizer.to_json()

      加载——keras.preprocessing.text.tokenizer_from_json

      在to_json()方法中,调用“get_config”方法来处理:

          json_word_counts = json.dumps(self.word_counts)
          json_word_docs = json.dumps(self.word_docs)
          json_index_docs = json.dumps(self.index_docs)
          json_word_index = json.dumps(self.word_index)
          json_index_word = json.dumps(self.index_word)
      
          return {
              'num_words': self.num_words,
              'filters': self.filters,
              'lower': self.lower,
              'split': self.split,
              'char_level': self.char_level,
              'oov_token': self.oov_token,
              'document_count': self.document_count,
              'word_counts': json_word_counts,
              'word_docs': json_word_docs,
              'index_docs': json_index_docs,
              'index_word': json_index_word,
              'word_index': json_word_index
          }
      

      【讨论】:

      • 正如目前所写,您的答案尚不清楚。请edit 添加其他详细信息,以帮助其他人了解这如何解决所提出的问题。你可以找到更多关于如何写好答案的信息in the help center。
      猜你喜欢
      • 2018-08-20
      • 1970-01-01
      • 1970-01-01
      • 2021-11-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多