【发布时间】:2019-01-14 12:01:48
【问题描述】:
我想了解创建句子的 fastText 向量的方式。根据这个issue 309,句子的向量是通过对单词的向量进行平均得到的。
为了确认这一点,我编写了以下脚本:
import numpy as np
import fastText as ft
# Loading model for Finnish.
model = ft.load_model('cc.fi.300.bin')
# Getting word vectors for 'one' and 'two'.
one = model.get_word_vector('yksi')
two = model.get_word_vector('kaksi')
# Getting the sentence vector for the sentence "one two" in Finnish.
one_two = model.get_sentence_vector('yksi kaksi')
one_two_avg = (one + two) / 2
# Checking if the two approaches yield the same result.
is_equal = np.array_equal(one_two, one_two_avg)
# Printing the result.
print(is_equal)
# Result: FALSE
但是,似乎获得的向量并不相似。
为什么两个值不一样?它会与我平均向量的方式有关吗?或者,也许我错过了什么?
【问题讨论】: