模型

Qwen3 235B-A22B

Qwen/Qwen3-235B-A22B

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:已人工审核 许可证:Apache License 2.0 上游版本:8efa61729e24

中文简介

Qwen3 235B-A22B 是面向集群和企业级评测的超大规模混合专家模型,总参数约 235B、激活参数约 22B。它适合具备专业基础设施的团队,不应按普通本地模型理解其存储、部署与运维要求。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Qwen3 235B-A22B 是超大规模混合专家模型。官方模型卡列出约 235B 总参数和 22B 激活参数,适合由具备专业多 GPU 或集群基础设施的团队评估。

模型原生上下文长度为 32,768 tokens,可按官方 YaRN 说明扩展至最高 131,072 tokens,并支持思考与非思考模式。

完整权重规模、设备互联、并行策略和运维能力都会影响是否可用。激活参数不能替代对磁盘、内存、传输和框架兼容性的完整评估。

译文已人工复核 · 本站基于上游 README 整理并人工复核 · 2026-07-23 23:47

Qwen3-235B-A22B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-235B-A22B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non-Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 32,768 natively and 131,072 tokens with YaRN. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. Quicksta

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合大型组织开展高能力模型评测、复杂推理与代码任务研究,以及已有多 GPU 或集群环境中的推理验证。普通个人设备通常不适合作为首选验证环境。

模型参数

混合专家模型,总参数约 235B、激活参数约 22B;原生上下文长度为 32,768 tokens,可按 YaRN 说明扩展至最高 131,072 tokens。支持思考与非思考模式。

文件说明

已审核版本的 Hugging Face 官方 API 返回 128 个文件条目,总存储量约 437.91 GB。这是完整仓库元数据规模,交付前需进一步明确文件范围、存储介质和校验方案。

上游文件元数据

  • .gitattributes1.53 KB
  • config.json965 B
  • generation_config.json239 B
  • LICENSE11.08 KB
  • merges.txt1.59 MB
  • model-00001-of-00118.safetensors3.72 GB
  • model-00002-of-00118.safetensors3.72 GB
  • model-00003-of-00118.safetensors3.72 GB
  • model-00004-of-00118.safetensors3.71 GB
  • model-00005-of-00118.safetensors3.72 GB
  • model-00006-of-00118.safetensors3.72 GB
  • model-00007-of-00118.safetensors3.72 GB

硬件建议

应按多 GPU 或集群工作负载设计,重点评估完整权重存储、精度、并行策略、设备互联、KV Cache、加载时间和故障恢复。本站不建议在缺少专业环境时直接以此档起步。

注意事项

该模型需要专业容量规划、框架适配和运维保障。激活参数不等于完整权重规模;数据安全、输出审核、系统权限、供应链和推理成本也必须纳入评估。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。