中文简介
这是基于 Qwen3.6-27B 的 1-bit Bonsai GGUF 版本,面向 llama.cpp 的 CUDA、Metal 和 CPU 推理。上游称模型部署体积约 3.9GB,并给出保留推理与智能体能力的评测结果;相关体积、速度和质量为上游测试结果,需在目标硬件上重新验证。
上游模型卡 / 数据集卡
这是基于 Qwen3.6-27B 的 1-bit Bonsai GGUF 版本,面向 llama.cpp 的 CUDA、Metal 和 CPU 推理。上游称模型部署体积约 3.9GB,并给出保留推理与智能体能力的评测结果;相关体积、速度和质量为上游测试结果,需在目标硬件上重新验证。
Prism ML Website | Whitepaper | Demo & Examples | Discord 1-bit Bonsai 27B — GGUF Full 27B-class reasoning in binary transformer weights, for llama.cpp (CUDA, Metal, CPU) **\~14.2x** smaller than FP16 | **\~90%** of FP16 intelligence retained | **\~44 tok/s** on an Apple M5 Pro laptop Highlights **\~3.9 GB** deployed footprint (down from \~54 GB FP16) — a 27B model on everyday laptops and single GPUs **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 **End-to-end binary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization **GGUF Q1_0_g128** format with custom 1-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16 **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.37x** decode speedup on the CUDA serving path **MLX companion**: also available as Bonsai-27B-mlx-1bit for native Apple Silicon inference, including iPhone (\~11 tok/s on iPhone 17 Pro Max via MLX Swift) **Ternary companion**: the quality-oriented operating point (\~7.2 GB, 95% of FP16) is also published in GGUF as Ternary-Bonsai-27B-gguf Resources **Whitepaper** — full methodology, benchmarks, and measurement notes **Demo & examples** — serving, benchmarking, and integrating Bonsai **Low-bit kernels**: llama.cpp fork (CUDA + Metal) · MLX fork (Apple Silicon) · mlx-swift fork (iOS/macOS) **Discord** — join the community fo
上游文件元数据
.eval_results/aime_2026.yaml175 B.eval_results/gsm8k.yaml153 B.eval_results/mmmu_pro.yaml165 B.gitattributes1.91 KBassets/bonsai-logo.svg515 BBonsai-27B-dspark-bf16.gguf6.79 GBBonsai-27B-dspark-Q4_1.gguf1.66 GBBonsai-27B-F16.gguf50.11 GBBonsai-27B-mmproj-BF16.gguf888.01 MBBonsai-27B-mmproj-Q8_0.gguf600.10 MBBonsai-27B-Q1_0.gguf3.54 GBLICENSE.txt9.94 KB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。