【问题标题】:How to manually calculate TF-IDF score from SKLearn's TfidfVectorizer如何从 SKLearn 的 TfidfVectorizer 手动计算 TF-IDF 分数
【发布时间】:2020-06-06 04:32:00
【问题描述】:

我一直在运行 SKLearn 的 TF-IDF Vectorizer,但在手动重新创建值时遇到了麻烦(作为帮助理解正在发生的事情)。

为了添加一些上下文,我有一个文档列表,我从中提取了命名实体(在我的实际数据中,这些文档最多为 5-gram,但在这里我将其限制为 bigrams)。我只想知道这些值的 TF-IDF 分数,并认为通过 vocabulary 参数传递这些术语可以做到这一点。

这是一些类似于我正在使用的虚拟数据:

from sklearn.feature_extraction.text import TfidfVectorizer
import pandas as pd    


# list of named entities I want to generate TF-IDF scores for
named_ents = ['boston','america','france','paris','san francisco']

# my list of documents
docs = ['i have never been to boston',
    'boston is in america',
    'paris is the capitol city of france',
    'this sentence has no named entities included',
    'i have been to san francisco and paris']

# find the max nGram in the named entity vocabulary
ne_vocab_split = [len(word.split()) for word in named_ents]
max_ngram = max(ne_vocab_split)

tfidf = TfidfVectorizer(vocabulary = named_ents, stop_words = None, ngram_range=(1,max_ngram))
tfidf_vector = tfidf.fit_transform(docs)

output = pd.DataFrame(tfidf_vector.T.todense(), index=named_ents, columns=docs)

注意:我知道默认情况下会删除停用词,但我的实际数据集中的一些命名实体包含诸如“国务院”之类的短语。所以他们一直留在这里。

这是我需要帮助的地方。我的理解是,我们计算 TF-IDF 如下:

TF: 词频:根据SKlearn guidelines 是“一个词在给定文档中出现的次数”

IDF: 逆文档频率:1+文档数与1+包含词条的文档数之比的自然对数。根据链接中的相同准则,结果值添加了 1 以防止被零除。

然后,我们将 TF 乘以 IDF,得到给定文档中给定术语的整体 TF-IDF。

示例

我们以第一列为例,它只有一个命名实体“波士顿”,根据上面的代码,在 1 的第一个文档上有一个 TF-IDF。但是,当我手动解决这个问题时,我得到以下:

TF = 1

IDF = log-e(1+total docs / 1+docs with 'boston') + 1
' ' = log-e(1+5 / 1+2) + 1
' ' = log-e(6 / 3) + 1
' ' = log-e(2) + 1
' ' = 0.69314 + 1
' ' = 1.69314

TF-IDF = 1 * 1.69314 = 1.69314 (not 1)

也许我在文档中遗漏了一些内容,即分数上限为 1,但我无法弄清楚我哪里出错了。此外,通过上述计算,第一列和第二列中波士顿的得分应该没有任何差异,因为该术语在每个文档中只出现一次。

编辑 在发布问题后,我认为术语频率可能被计算为与文档中单字元数或文档中命名实体数的比率。例如,在第二个文档中,SKlearn 为波士顿生成了0.627914 的分数。如果我将 TF 计算为令牌的比率 = 'boston' (1) : all unigram tokens (4) 我得到一个 0.25 的 TF,当我应用到 TF-IDF 时返回的分数刚好超过 0.147 .

类似地,当我使用令牌的比率 = 'boston' (1) : 所有 NE 令牌 (2) 并应用 TF-IDF 时,我得到了 0.846 的分数。很明显我在某个地方出错了。

【问题讨论】:

  • 一个可能有帮助的方法是逐步分解它。例如,来自 tfidfvectorizer 文档:“等效于 CountVectorizer,后跟 TfidfTransformer。”您可以尝试重新创建这些步骤并比较中间输出
  • 您缺少标准化步骤。请参阅下面的示例。

标签: python scikit-learn tf-idf tfidfvectorizer


【解决方案1】:

让我们一步一步地做这个数学练习。

第 1 步。获取 boston 令牌的 tfidf 分数

docs = ['i have never been to boston',
        'boston is in america',
        'paris is the capitol city of france',
        'this sentence has no named entities included',
        'i have been to san francisco and paris']

from sklearn.feature_extraction.text import TfidfVectorizer

# I did not include your named_ents here but did for a full vocab 
tfidf = TfidfVectorizer(smooth_idf=True,norm='l1')

注意TfidfVectorizer中的参数,它们对于以后的平滑和标准化很重要。

docs_tfidf = tfidf.fit_transform(docs).todense()
n = tfidf.vocabulary_["boston"]
docs_tfidf[:,n]
matrix([[0.19085885],
        [0.22326669],
        [0.        ],
        [0.        ],
        [0.        ]])

到目前为止,我们得到了 boston 令牌的 tfidf 分数(词汇中的第 3 名)。

步骤 2.计算 boston 令牌 w/o norm 的 tfidf。

公式是:

tf-idf(t, d) = tf(t, d) * idf(t)
idf(t) = log( (n+1) / (df(t)+1) ) + 1
其中:
- tf(t,d) -- 文档 d 中的简单词条 t 频率
- idf(t) -- 平滑反转文档频率(因为smooth_idf=True 参数)

计算第 0 个文档中的令牌 boston 和它出现在的文档数:

tfidf_boston_wo_norm = ((1/5) * (np.log((1+5)/(1+2))+1))
tfidf_boston_wo_norm
0.3386294361119891

注意,根据内置标记化方案,i 不算作标记。

第 3 步。标准化

让我们先进行l1 归一化,即所有计算的非归一化 tfdid 的总和应为 1:

l1_norm = ((1/5) * (np.log((1+5)/(1+2))+1) +
         (1/5) * (np.log((1+5)/(1+1))+1) +
         (1/5) * (np.log((1+5)/(1+2))+1) +
         (1/5) * (np.log((1+5)/(1+2))+1) +
         (1/5) * (np.log((1+5)/(1+2))+1))
tfidf_boston_w_l1_norm = tfidf_boston_wo_norm/l1_norm
tfidf_boston_w_l1_norm 
0.19085884520912985

如您所见,我们得到的 tfidf 分数与上述相同。

现在让我们对l2 norm 做同样的数学运算。

基准测试:

tfidf = TfidfVectorizer(sublinear_tf=True,norm='l2')
docs_tfidf = tfidf.fit_transform(docs).todense()
docs_tfidf[:,n]
matrix([[0.42500138],
        [0.44400208],
        [0.        ],
        [0.        ],
        [0.        ]])

微积分:

l2_norm = np.sqrt(((1/5) * (np.log((1+5)/(1+2))+1))**2 +
                  ((1/5) * (np.log((1+5)/(1+1))+1))**2 +
                  ((1/5) * (np.log((1+5)/(1+2))+1))**2 +
                  ((1/5) * (np.log((1+5)/(1+2))+1))**2 +
                  ((1/5) * (np.log((1+5)/(1+2))+1))**2                
                 )

tfidf_boston_w_l2_norm = tfidf_boston_wo_norm/l2_norm
tfidf_boston_w_l2_norm 
0.42500137513291814

还是和你看到的一样。

【讨论】:

    猜你喜欢
    • 2019-08-13
    • 2019-01-10
    • 2018-07-11
    • 2016-09-07
    • 2021-02-13
    • 2012-04-23
    • 2018-11-09
    • 2017-04-21
    • 2017-03-07
    相关资源
    最近更新 更多