模型

Laguna-S-2.1-NVFP4

poolside/Laguna-S-2.1-NVFP4

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:openmdw-1.1 上游版本:07614121b318

中文简介

Laguna S 2.1-NVFP4 是 Laguna S 2.1 的 NVFP4 压缩版本,面向本地智能体编程和长流程任务。上游说明其总参数约 117.6B、每 token 激活约 8.5B,并采用滑动窗口与全局注意力混合结构;硬件兼容性、精度损失和 KV 缓存占用需在目标推理框架中验证。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Laguna S 2.1-NVFP4 是 Laguna S 2.1 的 NVFP4 压缩版本,面向本地智能体编程和长流程任务。上游说明其总参数约 117.6B、每 token 激活约 8.5B,并采用滑动窗口与全局注意力混合结构;硬件兼容性、精度损失和 KV 缓存占用需在目标推理框架中验证。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Use on OpenRouter · Use on Vercel AI Gateway · Release blog post Laguna S 2.1-NVFP4 Laguna S 2.1-NVFP4 is a 117.6B total parameter Mixture-of-Experts model with 8.5B activated parameters per token designed for agentic coding and long-horizon work on a local machine. It uses Sliding Window Attention with per-head gating in 36 out of 48 layers for fast inference and low KV cache requirements. Highlights **Mixed SWA and global attention layout**: Laguna S 2.1 uses softplus gating with per-layer rotary scales, enabling mixed SWA (Sliding Window Attention) and global attention layers in a 3:1 ratio (across 48 total layers) **KV cache in FP8**: KV cache quantized to FP8, reducing memory per token **Native reasoning support**: Interleaved thinking between tool calls with support for enabling and disabling thinking per-request **Local-ready**: At 117.6B total parameters and 8.5B activated, the NVFP4 weights are roughly 71 GB. Available on Ollama and llama.cpp (BF16 and Q4\_K\_M only) **OpenMDW-1.1 license**: Use and modify the model and associated materials freely for commercial and non-commercial purposes (learn more about OpenMDW) Model overview Training: pre-training, post-training and reinforcement learning stages Number of parameters: 117.6B total with 8.5B activated per token Optimizer: Muon Layers: 48 layers (12 layers with global attention, 36 layers with sliding window attention) Experts: 256 experts with 1 shared expert Sliding Window: 512 tokens Modality: text-to-text Context window: 262,144 tokens Reasoning support: interleaved thinking with preserved thinking Benchmark results | Model | Size | Terminal-Bench 2.1 | SWE-bench Multilingual | SWE-Bench Pro (Public Dataset) | DeepSWE | SWE Atlas (Codebase QnA) | Toolathlon Verified | |---|---|---|---|---|---|---|---| | **Laguna S 2.1** | 118B-A8B | **70.2%** | **78.5%** | **59.4%** | **40.4%** | **46.2%** | **49.7%** | | Tencent Hy3 | 295B-A21B | 71.7% | 75.8% | 57.9% | - | - | - | | Inkling | 975B-A41B | 63.8% | -

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes1.55 KB
  • chat_template.jinja3.87 KB
  • config.json6.76 KB
  • configuration_laguna.py13.04 KB
  • generation_config.json485 B
  • LICENSE.md2.56 KB
  • model-00001-of-00015.safetensors4.76 GB
  • model-00002-of-00015.safetensors4.77 GB
  • model-00003-of-00015.safetensors4.77 GB
  • model-00004-of-00015.safetensors4.77 GB
  • model-00005-of-00015.safetensors4.77 GB
  • model-00006-of-00015.safetensors4.77 GB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。