【问题标题】:Using Bloom AI Model on Mac M1 for continuing prompts (Pytorch)在 Mac M1 上使用 Bloom AI 模型进行持续提示 (Pytorch)
【发布时间】:2022-11-15 01:14:07
【问题描述】:

我尝试在我的 Macbook M1 Max 64GB 上运行 bigscience Bloom AI 模型,新安装的适用于 Mac M1 芯片的 pytorch 和运行的 Python 3.10.6。 我根本无法获得任何输出。 对于其他 AI 模型,我也遇到了同样的问题,我真的不知道该如何解决。

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

device = "mps" if torch.backends.mps.is_available() else "cpu"
if device == "cpu" and torch.cuda.is_available():
    device = "cuda" #if the device is cpu and cuda is available, set the device to cuda
print(f"Using {device} device") #print the device

tokenizer = AutoTokenizer.from_pretrained("bigscience/bloom")
model = AutoModelForCausalLM.from_pretrained("bigscience/bloom").to(device)

input_text = "translate English to German: How old are you?"
input_ids = tokenizer(input_text, return_tensors="pt").input_ids.to(device)

outputs = model.generate(input_ids)
print(tokenizer.decode(outputs[0]))

我已经尝试过使用其他模型(较小的 bert 模型),还尝试让它在 CPU 上运行而不使用 mps 设备。

也许任何人都可以提供帮助

【问题讨论】:

  • 如果它很重要:我使用的是 113.0 Beta (22A5352e),但我想这应该不是问题

标签: machine-learning pytorch apple-m1 metal huggingface-transformers


【解决方案1】:

获得输出可能需要很长时间。您想将其分解为涉及的串行调用吗? a) 嵌入层 b) 70 个布隆块 c) 然后是输出层规范和 d) 令牌解码?

https://nbviewer.org/urls/arteagac.github.io/blog/bloom_local.ipynb 提供了运行此代码的示例。

它基本上归结为:

def forward(input_ids):
    # 1. Create attention mask and position encodings
    attention_mask = torch.ones(len(input_ids)).unsqueeze(0).bfloat16().to(device)
    alibi = build_alibi_tensor(input_ids.shape[1], config.num_attention_heads,
                               torch.bfloat16).to(device)
    # 2. Load and use word embeddings
    embeddings, lnorm = load_embeddings()
    hidden_states = lnorm(embeddings(input_ids))
    del embeddings, lnorm

    # 3. Load and use the BLOOM blocks sequentially
    for block_num in range(70):
        load_block(block, block_num)
        hidden_states = block(hidden_states, attention_mask=attention_mask, alibi=alibi)[0]
        print(".", end='')
    
    hidden_states = final_lnorm(hidden_states)
    
    #4. Load and use language model head
    lm_head = load_causal_lm_head()
    logits = lm_head(hidden_states)

    # 5. Compute next token 
    return torch.argmax(logits[:, -1, :], dim=-1)

请参阅链接的笔记本以获取forward 调用中使用的函数的实现。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2020-11-10
    • 2021-07-26
    • 1970-01-01
    • 2021-04-02
    • 2022-11-06
    • 2021-08-12
    • 2021-11-03
    • 2021-07-01
    相关资源
    最近更新 更多