【问题标题】:understanding top n tfidf features in TfidfVectorizer了解 TfidfVectorizer 中的前 n 个 tfidf 功能
【发布时间】:2020-05-07 05:36:10
【问题描述】:

我试图更好地理解scikit-learnTfidfVectorizer。下面的代码有两个文档doc1 = The car is driven on the road,doc2 = The truck is driven on the highway。通过调用fit_transform 会生成一个 tf-idf 权重的向量化矩阵。

根据tf-idf 值矩阵,highway,truck,car 不应该是最重要的词,而不是 highway,truck,driven 作为highway = truck= car= 0.63 and driven = 0.44

#testing tfidfvectorizer
from sklearn.feature_extraction.text import TfidfVectorizer
import numpy as np

tn = ['The car is driven on the road', 'The truck is driven on the highway']
vectorizer = TfidfVectorizer(tokenizer= lambda x:x.split(),stop_words = 'english')
response = vectorizer.fit_transform(tn)

feature_array = np.array(vectorizer.get_feature_names()) #list of features
print(feature_array)
print(response.toarray())

sorted_features = np.argsort(response.toarray()).flatten()[:-1] #index of highest valued features
print(sorted_features)

#printing top 3 weighted features
n = 3
top_n = feature_array[sorted_features][:n]
print(top_n)
['car' 'driven' 'highway' 'road' 'truck']
[[0.6316672  0.44943642 0.         0.6316672  0.        ]
 [0.         0.44943642 0.6316672  0.         0.6316672 ]]
[2 4 1 0 3 0 3 1 2]
['highway' 'truck' 'driven']

【问题讨论】:

    标签: python scikit-learn tf-idf tfidfvectorizer


    【解决方案1】:

    从结果可以看出,tf-idf 矩阵确实给highway,truck,car(和truck)的分数更高:

    tn = ['The car is driven on the road', 'The truck is driven on the highway']
    vectorizer = TfidfVectorizer(stop_words = 'english')
    response = vectorizer.fit_transform(tn)
    terms = vectorizer.get_feature_names()
    
    pd.DataFrame(response.toarray(), columns=terms)
    
            car    driven   highway      road     truck
    0  0.631667  0.449436  0.000000  0.631667  0.000000
    1  0.000000  0.449436  0.631667  0.000000  0.631667
    

    问题在于您通过展平阵列进行的进一步检查。要获得所有行的最高分,您可以改为执行以下操作:

    max_scores = response.toarray().max(0).argsort()
    np.array(terms)[max_scores[-4:]]
    array(['car', 'highway', 'road', 'truck'], dtype='<U7')
    

    其中得分最高的是在数据框中具有0.63 分数的特征名称。

    【讨论】:

    • max(0) response.toarray().max(0) 中是什么意思?
    • 您将获得每列的最大值。那就是轴参数ndarray.max(axis=0)@afsara_ben
    猜你喜欢
    • 2018-11-22
    • 2021-05-26
    • 2019-03-29
    • 1970-01-01
    • 2011-04-21
    • 2015-04-17
    • 1970-01-01
    • 2018-05-24
    • 2014-08-06
    相关资源
    最近更新 更多