跳到正文
原文
Hugging Face Blog·· 10 天前

Transformers 支持运行 llama.cpp GGUF 量化模型

Transformers now runs llama.cpp quants

SI 导读

Hugging Face 宣布 transformers 库新增对 GGUF 格式模型的支持,允许用户通过 from_pretrained 接口在本地加载并运行量化模型。该功能目前主要面向 Apple Silicon 设备,通过复用 ggml 内核实现接近 llama.cpp 的推理性能,并支持通过 transformers serve 暴露 OpenAI 兼容 API。

精选SI 评分82
推荐理由

原文展示了在 transformers 中直接加载 GGUF 量化模型的方法,并给出了 Apple Silicon 上的性能对比数据。

来源:Hugging Face Blog · huggingface.co

© 2026 SI·Hot · Super Intelligence Hot · 超级智能热点 · 网站数据均来源于网络公开资料,版权归来源方所有