中文简介
这是基于 Qwen3.6-27B 的三值权重 GGUF 版本,面向 llama.cpp、CUDA、Metal 和 CPU 等本地推理环境。上游称其部署体积约 7.2GB,并给出相对 FP16 的压缩与评测结果;这些性能数字属于上游声明,实际速度、质量和内存占用取决于设备、运行参数与软件版本。
上游模型卡 / 数据集卡
这是基于 Qwen3.6-27B 的三值权重 GGUF 版本,面向 llama.cpp、CUDA、Metal 和 CPU 等本地推理环境。上游称其部署体积约 7.2GB,并给出相对 FP16 的压缩与评测结果;这些性能数字属于上游声明,实际速度、质量和内存占用取决于设备、运行参数与软件版本。
Prism ML Website | Whitepaper | Demo & Examples | Discord Ternary Bonsai 27B — GGUF Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) **\~9.4x** smaller than FP16 (ideal) | **95%** of FP16 intelligence retained | **\~26 tok/s** on an Apple M5 Pro laptop Highlights **\~7.2 GB** deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU **95% of FP16 intelligence retained**: 80.49 average across 15 thinking-mode benchmarks — a *higher* score than the conventional IQ2_XXS build (72.73) at less than two-thirds of its footprint **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 **End-to-end ternary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.71 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization **GGUF Q2_0_g128** format with custom 2-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16 **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.34x** decode speedup on the CUDA serving path **MLX companion**: also available as Ternary-Bonsai-27B-mlx-2bit for native Apple Silicon inference **1-bit companion**: the phone-class operating point (\~3.9 GB) that fits an iPhone 17 Pro Max, published in GGUF as Bonsai-27B-gguf Resources **Whitepaper** — full methodology, benchmarks, and measurement notes **Demo & examples** — serving, benchmarking, and integrating Bonsai **Low-bi
上游文件元数据
.eval_results/aime_2026.yaml182 B.eval_results/gsm8k.yaml162 B.eval_results/mmmu_pro.yaml173 B.gitattributes2.09 KBassets/bonsai-logo.svg515 BLICENSE.txt9.94 KBNOTICE.txt411 BREADME.md22.55 KBTernary-Bonsai-27B-dspark-bf16.gguf6.79 GBTernary-Bonsai-27B-dspark-Q4_1.gguf1.81 GBTernary-Bonsai-27B-F16.gguf50.11 GBTernary-Bonsai-27B-mmproj-BF16.gguf888.01 MB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。