中文简介
Qwen3 235B-A22B 是面向集群和企业级评测的超大规模混合专家模型,总参数约 235B、激活参数约 22B。它适合具备专业基础设施的团队,不应按普通本地模型理解其存储、部署与运维要求。
上游模型卡 / 数据集卡
Qwen3 235B-A22B 是超大规模混合专家模型。官方模型卡列出约 235B 总参数和 22B 激活参数,适合由具备专业多 GPU 或集群基础设施的团队评估。
模型原生上下文长度为 32,768 tokens,可按官方 YaRN 说明扩展至最高 131,072 tokens,并支持思考与非思考模式。
完整权重规模、设备互联、并行策略和运维能力都会影响是否可用。激活参数不能替代对磁盘、内存、传输和框架兼容性的完整评估。
Qwen3-235B-A22B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-235B-A22B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non-Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 32,768 natively and 131,072 tokens with YaRN. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. Quicksta
适用场景
适合大型组织开展高能力模型评测、复杂推理与代码任务研究,以及已有多 GPU 或集群环境中的推理验证。普通个人设备通常不适合作为首选验证环境。
模型参数
混合专家模型,总参数约 235B、激活参数约 22B;原生上下文长度为 32,768 tokens,可按 YaRN 说明扩展至最高 131,072 tokens。支持思考与非思考模式。
文件说明
已审核版本的 Hugging Face 官方 API 返回 128 个文件条目,总存储量约 437.91 GB。这是完整仓库元数据规模,交付前需进一步明确文件范围、存储介质和校验方案。
上游文件元数据
.gitattributes1.53 KBconfig.json965 Bgeneration_config.json239 BLICENSE11.08 KBmerges.txt1.59 MBmodel-00001-of-00118.safetensors3.72 GBmodel-00002-of-00118.safetensors3.72 GBmodel-00003-of-00118.safetensors3.72 GBmodel-00004-of-00118.safetensors3.71 GBmodel-00005-of-00118.safetensors3.72 GBmodel-00006-of-00118.safetensors3.72 GBmodel-00007-of-00118.safetensors3.72 GB
硬件建议
应按多 GPU 或集群工作负载设计,重点评估完整权重存储、精度、并行策略、设备互联、KV Cache、加载时间和故障恢复。本站不建议在缺少专业环境时直接以此档起步。
注意事项
该模型需要专业容量规划、框架适配和运维保障。激活参数不等于完整权重规模;数据安全、输出审核、系统权限、供应链和推理成本也必须纳入评估。
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。