模型

Bonsai-27B-mlx-1bit

prism-ml/Bonsai-27B-mlx-1bit

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:apache-2.0 上游版本:ef22f239c670

中文简介

这是面向 Apple MLX 生态的 1-bit Bonsai 27B 版本。上游强调约 3.9GB 的部署体积和移动设备运行能力,并给出相对 FP16 的质量与速度测试;手机可运行性、内存限制、散热和实际吞吐均高度依赖具体设备与软件版本。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

这是面向 Apple MLX 生态的 1-bit Bonsai 27B 版本。上游强调约 3.9GB 的部署体积和移动设备运行能力,并给出相对 FP16 的质量与速度测试;手机可运行性、内存限制、散热和实际吞吐均高度依赖具体设备与软件版本。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Prism ML Website  |  Whitepaper  |  Demo & Examples  |  Discord 1-bit Bonsai 27B Full 27B-class reasoning in binary transformer weights — the first 27B-class model to run on a phone **~14.2x** smaller than FP16 | **~90%** of FP16 intelligence retained | **~11 tok/s** on iPhone 17 Pro Max Highlights **~3.9 GB** deployed footprint (down from ~54 GB FP16) — fits within the per-app memory budget of a high-end phone such as the iPhone 17 Pro Max **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 **End-to-end binary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (~75% linear attention) and 4-bit KV-cache quantization **First interactive 27B-class generation on a phone**: ~11 tok/s on iPhone 17 Pro Max; ~44 tok/s on an Apple M5 Pro laptop **Custom 1-bit hybrid-attention kernels** on Apple MLX (Python, Swift) and CUDA — packed weights are consumed directly, never expanded back to FP16 **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.37x** decode speedup on the CUDA serving path **Ternary companion**: also available as Ternary Bonsai 27B, the quality-oriented operating point (~7.2 GB, 95% of FP16) for laptops and GPUs Resources **Whitepaper** — full methodology, benchmarks, and measurement notes **Demo & examples** — serving, benchmarking, and integrating Bonsai **Low-bit kernels**: MLX fork (Apple Silicon) · mlx-swift fork (iOS/macOS) · llama.cpp fork (CUDA) **Discord** — join the community for support, discussion

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes1.53 KB
  • assets/bonsai-logo.svg515 B
  • chat_template.jinja7.58 KB
  • config.json3.70 KB
  • LICENSE.txt9.94 KB
  • merges.txt3.20 MB
  • model.safetensors4.78 GB
  • model.safetensors.index.json181.23 KB
  • NOTICE.txt411 B
  • preprocessor_config.json390 B
  • processor_config.json44 B
  • README.md22.47 KB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。