模型

DeepSeek V4 Pro

deepseek-ai/DeepSeek-V4-Pro

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:AI 辅助整理并检查 许可证:MIT License 上游版本:b5968e9190ef

中文简介

DeepSeek V4 Pro 是 DeepSeek 发布方公开的 V4 系列大型 MoE 模型。固定 README 标注 1.6T 总参数、49B 激活参数和 1M 上下文,当前固定仓库约 805.37 GB;适合已有多机推理与长上下文工程能力的团队,不是普通单卡下载后即可运行的模型。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

版本定位:DeepSeek V4 Pro 是 V4 预览系列的大型 MoE 版本。固定卡片给出 1.6T 总参数、49B 激活参数、1M 上下文和 FP4+FP8 混合精度;本站只记录发布方固定快照,不把其自测榜单当作独立保证。

部署关键点:仓库约 805.37 GB,含 64 个权重分片及专用 encoding、inference 文件。README 明确说明没有 Jinja chat template,获取与交付时必须把权重、索引、配置和编码脚本作为同一 SHA 的完整集合。

服务说明:橙子AI科技可在许可证和平台规则允许的前提下协助核对版本、完整文件清单、SHA256 与网盘或硬盘交付方案;一期官网不托管权重,也不代表 DeepSeek 官方渠道。

已有简体中文译文 · Codex 基于固定版本上游 README 编写;待人工复核 · 2026-08-07 23:24

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Technical Report 👁️ Introduction We present a preview version of **DeepSeek-V4** series, including two strong Mixture-of-Experts (MoE) language models — **DeepSeek-V4-Pro** with 1.6T parameters (49B activated) and **DeepSeek-V4-Flash** with 284B parameters (13B activated) — both supporting a context length of **one million tokens**. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: 1. **Hybrid Attention Architecture:** We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only **27% of single-token inference FLOPs** and **10% of KV cache** compared with DeepSeek-V3.2. 2. **Manifold-Constrained Hyper-Connections (mHC):** We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity. 3. **Muon Optimizer:** We employ the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than **32T** diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model. DeepSeek-V4-Pro-Max**, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, **DeepSeek-V4-Flash-M

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合企业或高校评估长上下文推理、复杂代码与智能体任务,并用于验证 V4 专用编码、权重转换和分布式推理链路。若项目仍处于需求验证阶段,应先通过 API 或更小模型确认任务指标,再决定是否获取完整权重。

模型参数

固定版本:b5968e9190ef611bbf34a7229255be88a0e937c1。固定 README:MoE、1.6T 总参数、49B 激活参数、1M 上下文,MoE 专家参数 FP4、其余多数参数 FP8;Hugging Face Safetensors 元数据参数量为 1,598,839,674,782,架构为 DeepseekV4ForCausalLM。

文件说明

固定 SHA 下共有 91 个文件,包含 64 个 Safetensors 权重分片、索引、配置、tokenizer、encoding、inference 与技术报告,Hugging Face used_storage 约 805.37 GB。交付必须固定 SHA 并保留索引和专用编码文件,不能只拿权重分片。

上游文件元数据

  • .gitattributes1.63 KB
  • assets/dsv4_performance.png976.91 KB
  • config.json1.79 KB
  • encoding/encoding_dsv4.py27.25 KB
  • encoding/README.md7.93 KB
  • encoding/test_encoding_dsv4.py3.65 KB
  • encoding/tests/test_input_1.json2.68 KB
  • encoding/tests/test_input_2.json526 B
  • encoding/tests/test_input_3.json4.44 KB
  • encoding/tests/test_input_4.json2.67 KB
  • encoding/tests/test_output_1.txt2.33 KB
  • encoding/tests/test_output_2.txt342 B

硬件建议

约 805 GB 的仓库规模通常需要多卡或多节点存储与推理方案;1M 上下文还会增加 KV cache 和运行时开销。激活 49B 只描述每 token 计算量,不代表完整权重只占 49B 规模,也不能据此承诺单机可运行。

注意事项

固定仓库采用 MIT License,但业务上线仍需审查模型输出、工具权限、隐私与行业要求。README 明确说明该版本未提供 Jinja 格式 chat template,必须保留并审查 encoding 与 inference 目录;发布方榜单不等于客户环境可复现结果。

获取、校验与交付

咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。