【问题标题】:How to save a scikit-learn classifier that utilizes a vectorizer, a pipeline and GridSearchV?如何保存使用矢量化器、管道和 GridSearchV 的 scikit-learn 分类器?
【发布时间】:2020-08-01 17:49:08
【问题描述】:

我使用以下步骤构建了一个情绪分类器:

load dataset with pandas

count = CountVectorizer()
bag = count.fit_transform(x)
bag.toarray()
tfidf = TfidfTransformer(use_idf=True, norm="l2",smooth_idf=True)
tfidf.fit_transform(bag).toarray()

from collections import Counter

vocab = Counter()
for text in x:
    for word in text.split(" "):
        vocab[word] += 1

import nltk
from nltk.corpus import stopwords
stop = stopwords.words('english')

vocab_reduced = Counter()
for w, c in vocab.items():
    if not w in stop:
        vocab_reduced[w]=c

def preprocessor(text):
    """ Return a cleaned version of text
    """
    # Remove HTML markup
    text = re.sub('<[^>]*>', '', text)
    # Save emoticons for later appending
    emoticons = re.findall('(?::|;|=)(?:-)?(?:\)|\(|D|P)', text)
    # Remove any non-word character and append the emoticons,
    # removing the nose character for standarization. Convert to lower case
    text = (re.sub('[\W]+', ' ', text.lower()) + ' ' + ' '.join(emoticons).replace('-', ''))
    
    return text

from nltk.stem import PorterStemmer

porter = PorterStemmer()

def tokenizer(text):
    return text.split()

def tokenizer_porter(text):
    return [porter.stem(word) for word in text.split()]

tfidf = TfidfVectorizer(strip_accents=None,
                        lowercase=False,
                        preprocessor=None)

param_grid = [{'vect__ngram_range': [(1, 1)],
               'vect__stop_words': [stop, None],
               'vect__tokenizer': [tokenizer, tokenizer_porter],
               'vect__preprocessor': [None, preprocessor],
               'vect__use_idf':[False],
               'vect__norm':[None],
               "clf__alpha":[0,1],
               "clf__fit_prior":[False,True]},
                ]
multi_tfidf = Pipeline([("vect", tfidf),
                       ( "clf", MultinomialNB())])
gs_multi_tfidf = GridSearchCV(multi_tfidf, param_grid,
                              scoring="accuracy",
                              cv=5,
                              verbose=1,
                              n_jobs=-1)
gs_multi_tfidf.fit(X_train,y_train)

我尝试使用 joblib 保存管道并同时保存分类器和管道,然后将其用于网站。但每次我尝试,它都不起作用。我要么得到:ValueError: not enough values to unpack (expected 2, got 1)(同时保存了管道和分类器),要么得到 TypeError: 'module' object is not callable,只使用分类器。

【问题讨论】:

    标签: python scikit-learn pickle text-classification


    【解决方案1】:

    请尝试使用以下内容。为什么不包括 CountVectorizer() 和 TfidfTransformer() 的任何具体原因?您还应该准确指定您尝试保存模型的方式。

    multi_tfidf = Pipeline([("vect", TfidfVectorizer()),
                           ( "clf", MultinomialNB())])
    

    【讨论】:

    • 我用了这个教程:github.com/Abhishek-MLDL/logistic-sentiment/blob/master/…,他也没有像你说的那样做。关于我如何保存它: joblib.dump(Pipeline, "Finished Models/MultinominalNB/MultinominalNB_pipeline.joblib") 和 joblib.dump(gs_multi_tfidf.best_estimator_, "Finished Models/MultinominalNB/MultinominalNB_Grid.joblib") 你认为我应该预处理我的事先数据然后使用tfidftransformer?然后在没有管道的情况下使用 GridsearchV?
    猜你喜欢
    • 2015-03-02
    • 2018-07-01
    • 2019-06-09
    • 2020-12-19
    • 1970-01-01
    • 1970-01-01
    • 2017-12-25
    • 1970-01-01
    • 2016-10-25
    相关资源
    最近更新 更多