Transformers 支持运行 llama.cpp GGUF 量化模型
Hugging Face 宣布 transformers 库新增对 GGUF 格式模型的支持,允许用户通过 from_pretrained 接口在本地加载并运行量化模型。该功能目前主要面向 Apple Silicon 设备,通过复用 ggml 内核实现接近 llama.cpp 的推理性能,并支持通过 transformers serve 暴露 OpenAI 兼容 API。
推荐理由:原文展示了在 transformers 中直接加载 GGUF 量化模型的方法,并给出了 Apple Silicon 上的性能对比数据。