【发布时间】:2012-08-20 13:38:00
【问题描述】:
我正在学习Part 1 和Part 2 上提供的教程。不幸的是,作者没有时间在最后一节中使用余弦相似度来实际找到两个文档之间的距离。我在stackoverflow的以下链接的帮助下按照文章中的示例进行操作,包括上面链接中提到的代码(只是为了让生活更轻松)
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.feature_extraction.text import TfidfTransformer
from nltk.corpus import stopwords
import numpy as np
import numpy.linalg as LA
train_set = ["The sky is blue.", "The sun is bright."] # Documents
test_set = ["The sun in the sky is bright."] # Query
stopWords = stopwords.words('english')
vectorizer = CountVectorizer(stop_words = stopWords)
#print vectorizer
transformer = TfidfTransformer()
#print transformer
trainVectorizerArray = vectorizer.fit_transform(train_set).toarray()
testVectorizerArray = vectorizer.transform(test_set).toarray()
print 'Fit Vectorizer to train set', trainVectorizerArray
print 'Transform Vectorizer to test set', testVectorizerArray
transformer.fit(trainVectorizerArray)
print
print transformer.transform(trainVectorizerArray).toarray()
transformer.fit(testVectorizerArray)
print
tfidf = transformer.transform(testVectorizerArray)
print tfidf.todense()
由于上面的代码,我有以下矩阵
Fit Vectorizer to train set [[1 0 1 0]
[0 1 0 1]]
Transform Vectorizer to test set [[0 1 1 1]]
[[ 0.70710678 0. 0.70710678 0. ]
[ 0. 0.70710678 0. 0.70710678]]
[[ 0. 0.57735027 0.57735027 0.57735027]]
我不确定如何使用此输出来计算余弦相似度,我知道如何针对两个长度相似的向量实现余弦相似度,但在这里我不确定如何识别这两个向量。
【问题讨论】:
-
对于trainVectorizerArray中的每个向量,你必须找到与testVectorizerArray中向量的余弦相似度。
-
@excray 谢谢你的帮助,我设法弄清楚了,我应该回答吗?
-
@excray 但我确实有一个小问题,实际上 tf*idf 计算对此没有用,因为我没有使用矩阵中显示的最终结果。
-
这是您引用的教程的第三部分,详细回答了您的问题pyevolve.sourceforge.net/wordpress/?p=2497
-
@ClémentRenaud 我点击了您提供的链接,但由于我的文档较大,它开始抛出 MemoryError 我们该如何处理?
标签: python machine-learning nltk information-retrieval tf-idf