【问题标题】:Get Bert Embeddings for every Token in a Sentence为句子中的每个标记获取 Bert 嵌入
【发布时间】:2021-05-31 23:56:13
【问题描述】:

我在 python 中有一个数据框,其中有一列文本数据。我需要运行一个循环,在该循环中,我将获取该文本列中的每一行,并为该特定行中的每个标记获取 bert 嵌入。然后我需要附加这些向量嵌入并出于某种目的进行尝试。

例如“我的名字是奥巴马” 为“我的”获得 768 个向量嵌入 得到 768 向量嵌入'name' 得到 768 向量嵌入 'is' 得到 'Obama' 的 768 个向量嵌入

最终输出:大小为 768*4 = 3072 的向量嵌入 假设每一行都有确切的单词数

【问题讨论】:

    标签: python pandas machine-learning nlp data-science


    【解决方案1】:

    我相信您正在尝试将句子中单个单词的基于上下文的嵌入引入图片,而不是像 GloVe 那样的固定向量。 你的方法应该是。

    1. 将您的段落标记为单个句子(如果适用,请查看一些句子标记器或 SBD(句子边界检测)方法)
    2. 现在对于构成段落的每个句子,获取单词的嵌入。
    3. 求平均值,以便您在多个段落(在您的情况下为数据框单元格 - 本质上是段落)中获得一致形状的向量

    pip install sentence-transformers

    一旦安装;

    model = SentenceTransformer('paraphrase-distilroberta-base-v1')
    
    #Our sentences we like to encode
    sentences = ['This framework generates embeddings for each input sentence',
        'Sentences are passed as a list of string.',
        'The quick brown fox jumps over the lazy dog.']
    
    #Sentences are encoded by calling model.encode()
    embeddings = model.encode(sentences)
    
    #Print the embeddings
    for sentence, embedding in zip(sentences, embeddings):
        print("Sentence:", sentence)
        print("Embedding:", embedding)
        print("")
    

    查看嵌入向量和围绕嵌入的聚合技术。

    【讨论】:

    • 这也将为每个句子提供 768 个向量嵌入。我希望为每个单词获得 768 个向量嵌入。有没有办法使用bert做到这一点?
    • 再看一下,这会给你 768 * n 向量。其中 n 是句子中的单词数。
    猜你喜欢
    • 2021-11-29
    • 2020-01-29
    • 1970-01-01
    • 1970-01-01
    • 2020-04-07
    • 1970-01-01
    • 2020-12-16
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多