中文简介
Qwen3 32B 是 Qwen3 系列的大型稠密模型,适合具备较充足计算资源、需要评估复杂推理与高质量生成的组织。与 MoE 档相比,它采用稠密架构,容量规划和性能表现应分别测试。
上游模型卡 / 数据集卡
Qwen3 32B 是大型稠密语言模型,支持 Qwen3 系列的思考与非思考双模式,适合对复杂推理、代码和知识工作进行高质量评估。
模型约 32.8B 参数,原生上下文长度为 32,768 tokens;上游模型卡给出 YaRN 扩展到最高 131,072 tokens 的说明。长上下文和高并发会进一步提高资源需求。
部署前应与较小模型和 MoE 模型进行同任务比较,并核对许可证、权重格式、推理框架与目标硬件。
Qwen3-32B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-32B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 32.8B Number of Paramaters (Non-Embedding): 31.2B Number of Layers: 64 Number of Attention Heads (GQA): 64 for Q and 8 for KV Context Length: 32,768 natively and 131,072 tokens with YaRN. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. Quickstart The code of Qwen3 has been in the latest Hugging Face `transformers` and we advise you t
适用场景
适合复杂代码、长文档分析、专业知识工作流和企业级质量评测。推荐先建立小规模基准,再比较 14B、30B-A3B 与 32B 的质量、时延和总体成本。
模型参数
约 32.8B 参数的稠密模型,原生上下文长度为 32,768 tokens;可按上游 YaRN 说明扩展至最高 131,072 tokens。支持思考和非思考模式。
文件说明
本次审核提交由 Hugging Face 官方 API 返回 27 个文件条目,总存储量约 61.03 GB。不同格式的文件范围和运行占用应在交付前逐项确认。
上游文件元数据
.gitattributes1.53 KBconfig.json728 Bgeneration_config.json239 BLICENSE11.08 KBmerges.txt1.59 MBmodel-00001-of-00017.safetensors3.69 GBmodel-00002-of-00017.safetensors3.63 GBmodel-00003-of-00017.safetensors3.63 GBmodel-00004-of-00017.safetensors3.63 GBmodel-00005-of-00017.safetensors3.63 GBmodel-00006-of-00017.safetensors3.63 GBmodel-00007-of-00017.safetensors3.63 GB
硬件建议
通常需要服务器级 GPU 或多设备方案,量化也不能替代真实容量验证。应测试框架兼容性、加载时间、显存峰值、上下文长度、并发吞吐和故障恢复。
注意事项
稠密 32B 对存储、内存和计算资源要求较高。输出仍可能出错,涉及生产决策、代码执行、法律、财务或安全内容时必须由专业人员复核。
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。