中文简介
DeepSeek R1 0528 完整版适合研究大型推理模型、多机推理和蒸馏流程,不是面向普通个人电脑的轻量下载项。本站核对的是发布方仓库固定版本 4236a6af538f,该版本模型元数据显示约 6845.31 亿参数;页面只提供版本、许可和交付咨询信息,不托管权重,也不代表 DeepSeek 官方服务。
上游模型卡 / 数据集卡
版本定位:这是 DeepSeek 发布方仓库中的 R1 0528 完整模型快照。固定 README 说明它是在 R1 基础上的小版本升级,重点面向推理、编程和工具使用改进;本站没有把发布方自测分数转换成独立排名。
选择建议:约 6845 亿参数和数百 GB 仓库存储决定了它更适合已经具备分布式推理条件的团队。个人或单卡用户通常应先评估同版本的 Qwen3 8B 蒸馏模型;二者用途和资源门槛不同,不能只按模型名称判断。
许可与交付:固定版本 LICENSE 为 MIT。本站可在权限和许可证允许的前提下协助核对 SHA、完整文件清单和校验结果;不代理模型在线推理,不承诺特定硬件一定可运行,也不替代客户对生成内容和应用场景的安全审查。
DeepSeek-R1-0528 Paper Link 👁️ 1. Introduction The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro. Compared to the previous version, the upgraded model shows significant improvements in handling complex reasoning tasks. For instance, in the AIME 2025 test, the model’s accuracy has increased from 70% in the previous version to 87.5% in the current version. This advancement stems from enhanced thinking depth during the reasoning process: in the AIME test set, the previous model used an average of 12K tokens per question, whereas the new version averages 23K tokens per question. Beyond its improved reasoning capabilities, this version also offers a reduced hallucination rate, enhanced support for function calling, and better experience for vibe coding. 2. Evaluation Results DeepSeek-R1-0528 For all our models, the maximum generation length is set to 64K tokens. For benchmarks requiring sampling, we use a temperature of $0.6$, a top-p value of $0.95$, and generate 16 responses per query to estimate pass@1. | Category | Benchmark (Metric) | DeepSeek R1 | DeepSeek R1 0528 |----------|----------------------------------|-----------------|---| | General | | | MMLU-Redux (EM) | 92.9 | 93.4 | | MMLU-Pro (EM) | 84.0 | 85.0 | | GPQA-Diamond (Pass@1) | 71.5 | 81.0 | | SimpleQA (Correct) | 30.1 | 27.8 | | FRAMES (Acc.) | 82.5 | 83.0 | | Humanity's Last Exam (Pass@1) | 8.5 | 17.7 | Code | | | LiveCodeBench (2408-2505) (Pass@1) | 63.5 | 73.3 | | Codeforces-Div
适用场景
适合具备多机多卡平台的企业或高校进行复杂推理评估、推理框架适配、吞吐与并行策略验证,以及为小模型蒸馏研究准备可追溯基线。若目标只是单机问答或课程实验,应优先比较同系列 8B 蒸馏版,避免先获取数百 GB 完整仓库后才发现部署条件不匹配。
模型参数
固定版本:4236a6af538feda4548eca9ab308586007567f52。Hugging Face Safetensors 元数据参数量:684,531,386,000;架构字段:DeepseekV3ForCausalLM;仓库标签包含 FP8。固定 README 说明其评测最长生成长度设置为 64K token,该数字不能直接当作所有推理框架的可用上下文承诺。
文件说明
Hugging Face API 快照记录 174 个文件,used_storage 字段约 641.30 GB。由于本站只保存前 100 条文件元数据,页面明确标注文件清单已截断;641.30 GB 是上游仓库存储字段,不等于某一种部署方案必须下载的精确字节数。交付前应按固定 SHA 重新导出完整文件清单与校验值。
上游文件元数据
.gitattributes1.52 KBconfig.json1.62 KBconfiguration_deepseek.py9.67 KBfigures/benchmark.png314.94 KBgeneration_config.json171 BLICENSE1.06 KBmodel-00001-of-000163.safetensors4.87 GBmodel-00002-of-000163.safetensors4.01 GBmodel-00003-of-000163.safetensors4.01 GBmodel-00004-of-000163.safetensors4.01 GBmodel-00005-of-000163.safetensors4.01 GBmodel-00006-of-000163.safetensors4.07 GB
硬件建议
这是多机多卡级资源。除权重空间外,还要为下载暂存、校验、解包或格式转换、运行时缓存和 KV cache 预留容量,落盘规划宜显著高于仓库标称规模。实际 GPU 数量取决于精度、并行框架、上下文和并发;不能仅按“激活参数”估算显存。普通工作站建议先验证 8B 蒸馏版。
注意事项
固定版本使用 MIT License,但许可证宽松不等于输出天然准确、合规或适合高风险决策。完整模型需要核对推理代码、分词器、精度格式和并行方案;发布方 README 中的基准成绩属于其测试条件,本站不作复现保证。任何转存或交付均以合法访问、上游规则和客户用途为前提。
获取、校验与交付
先核对上游版本、访问条件和许可证,再根据文件规模选择自行获取或人工协助。橙子AI科技可在合法访问权限、许可证及平台规则允许的前提下,提供版本与文件范围核对、下载协助、完整性校验和国内交付咨询;本官网本身不托管或下载资源文件。
- 大模型获取与国内转存服务
- 大模型怎么下载:Hugging Face 官方方式与国内交付完整指南
- 国内下载 Hugging Face 模型很慢或经常中断怎么办
- DeepSeek、ChatGPT、豆包、Kimi 能否下载到本地
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。