模型

Bonsai-27B-gguf

prism-ml/Bonsai-27B-gguf

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:apache-2.0 上游版本:f10afb355f10

中文简介

这是基于 Qwen3.6-27B 的 1-bit Bonsai GGUF 版本,面向 llama.cpp 的 CUDA、Metal 和 CPU 推理。上游称模型部署体积约 3.9GB,并给出保留推理与智能体能力的评测结果;相关体积、速度和质量为上游测试结果,需在目标硬件上重新验证。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

这是基于 Qwen3.6-27B 的 1-bit Bonsai GGUF 版本,面向 llama.cpp 的 CUDA、Metal 和 CPU 推理。上游称模型部署体积约 3.9GB,并给出保留推理与智能体能力的评测结果;相关体积、速度和质量为上游测试结果,需在目标硬件上重新验证。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Prism ML Website  |  Whitepaper  |  Demo & Examples  |  Discord 1-bit Bonsai 27B — GGUF Full 27B-class reasoning in binary transformer weights, for llama.cpp (CUDA, Metal, CPU) **\~14.2x** smaller than FP16 | **\~90%** of FP16 intelligence retained | **\~44 tok/s** on an Apple M5 Pro laptop Highlights **\~3.9 GB** deployed footprint (down from \~54 GB FP16) — a 27B model on everyday laptops and single GPUs **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 **End-to-end binary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization **GGUF Q1_0_g128** format with custom 1-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16 **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.37x** decode speedup on the CUDA serving path **MLX companion**: also available as Bonsai-27B-mlx-1bit for native Apple Silicon inference, including iPhone (\~11 tok/s on iPhone 17 Pro Max via MLX Swift) **Ternary companion**: the quality-oriented operating point (\~7.2 GB, 95% of FP16) is also published in GGUF as Ternary-Bonsai-27B-gguf Resources **Whitepaper** — full methodology, benchmarks, and measurement notes **Demo & examples** — serving, benchmarking, and integrating Bonsai **Low-bit kernels**: llama.cpp fork (CUDA + Metal) · MLX fork (Apple Silicon) · mlx-swift fork (iOS/macOS) **Discord** — join the community fo

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .eval_results/aime_2026.yaml175 B
  • .eval_results/gsm8k.yaml153 B
  • .eval_results/mmmu_pro.yaml165 B
  • .gitattributes1.91 KB
  • assets/bonsai-logo.svg515 B
  • Bonsai-27B-dspark-bf16.gguf6.79 GB
  • Bonsai-27B-dspark-Q4_1.gguf1.66 GB
  • Bonsai-27B-F16.gguf50.11 GB
  • Bonsai-27B-mmproj-BF16.gguf888.01 MB
  • Bonsai-27B-mmproj-Q8_0.gguf600.10 MB
  • Bonsai-27B-Q1_0.gguf3.54 GB
  • LICENSE.txt9.94 KB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。