Hugging Face Blog·· 7 天前精选AI 评分73
Hugging Face transformers 支持运行 llama.cpp 的 GGUF 量化模型
Transformers now runs llama.cpp quants
AI 导读
Hugging Face 在 transformers 中加入对 GGUF 模型的高效运行支持,通过复用 ggml 的 Metal 内核,可在 Apple Silicon 上以熟悉的 from_pretrained 接口加载并生成。
推荐理由
原文给出了 GGUF 加载方式、性能对比和当前限制,读者可以据此判断在 transformers 里跑本地量化模型的可行性与适用场景。
来源:Hugging Face Blog · huggingface.co