模型

Ternary-Bonsai-27B-gguf

prism-ml/Ternary-Bonsai-27B-gguf

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:apache-2.0 上游版本:abbae723028d

中文简介

这是基于 Qwen3.6-27B 的三值权重 GGUF 版本,面向 llama.cpp、CUDA、Metal 和 CPU 等本地推理环境。上游称其部署体积约 7.2GB,并给出相对 FP16 的压缩与评测结果;这些性能数字属于上游声明,实际速度、质量和内存占用取决于设备、运行参数与软件版本。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

这是基于 Qwen3.6-27B 的三值权重 GGUF 版本,面向 llama.cpp、CUDA、Metal 和 CPU 等本地推理环境。上游称其部署体积约 7.2GB,并给出相对 FP16 的压缩与评测结果;这些性能数字属于上游声明,实际速度、质量和内存占用取决于设备、运行参数与软件版本。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Prism ML Website  |  Whitepaper  |  Demo & Examples  |  Discord Ternary Bonsai 27B — GGUF Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) **\~9.4x** smaller than FP16 (ideal) | **95%** of FP16 intelligence retained | **\~26 tok/s** on an Apple M5 Pro laptop Highlights **\~7.2 GB** deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU **95% of FP16 intelligence retained**: 80.49 average across 15 thinking-mode benchmarks — a *higher* score than the conventional IQ2_XXS build (72.73) at less than two-thirds of its footprint **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 **End-to-end ternary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.71 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization **GGUF Q2_0_g128** format with custom 2-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16 **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.34x** decode speedup on the CUDA serving path **MLX companion**: also available as Ternary-Bonsai-27B-mlx-2bit for native Apple Silicon inference **1-bit companion**: the phone-class operating point (\~3.9 GB) that fits an iPhone 17 Pro Max, published in GGUF as Bonsai-27B-gguf Resources **Whitepaper** — full methodology, benchmarks, and measurement notes **Demo & examples** — serving, benchmarking, and integrating Bonsai **Low-bi

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .eval_results/aime_2026.yaml182 B
  • .eval_results/gsm8k.yaml162 B
  • .eval_results/mmmu_pro.yaml173 B
  • .gitattributes2.09 KB
  • assets/bonsai-logo.svg515 B
  • LICENSE.txt9.94 KB
  • NOTICE.txt411 B
  • README.md22.55 KB
  • Ternary-Bonsai-27B-dspark-bf16.gguf6.79 GB
  • Ternary-Bonsai-27B-dspark-Q4_1.gguf1.81 GB
  • Ternary-Bonsai-27B-F16.gguf50.11 GB
  • Ternary-Bonsai-27B-mmproj-BF16.gguf888.01 MB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。