模型

Qwen3 8B

Qwen/Qwen3-8B

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:已人工审核 许可证:Apache License 2.0 上游版本:b968826d9c46

中文简介

Qwen3 8B 是 Qwen3 系列中兼顾能力、部署复杂度与生态适配的通用稠密模型,支持思考和非思考模式。它适合企业原型、代码与知识工作评估,也是从轻量验证迈向较完整能力测试的常用档位。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Qwen3 8B 是 Qwen3 系列的通用稠密语言模型,官方模型卡强调思考模式与非思考模式的统一支持,可按任务选择更深入推理或更直接响应。

模型原生上下文长度为 32,768 tokens,并提供 YaRN 扩展到最高 131,072 tokens 的官方配置说明。超长上下文会影响资源、速度和质量,应根据实际材料长度进行评测。

它适合对话、代码和知识工作原型,但模型输出仍需事实核验与安全控制。许可证、部署示例和最新限制以上游固定版本页面为准。

译文已人工复核 · 本站基于上游 README 整理并人工复核 · 2026-07-23 23:47

Qwen3-8B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-8B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 8.2B Number of Paramaters (Non-Embedding): 6.95B Number of Layers: 36 Number of Attention Heads (GQA): 32 for Q and 8 for KV Context Length: 32,768 natively and 131,072 tokens with YaRN. For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. Quickstart The code of Qwen3 has been in the latest Hugging Face `transformers` and we advise you to u

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合通用对话、代码辅助、文档处理、知识库应用、工具调用原型和多语言内容生成。上线前应围绕准确性、提示注入、敏感信息和业务风险建立评测集。

模型参数

约 8.2B 参数的稠密模型,原生上下文长度为 32,768 tokens。上游说明可借助 YaRN 扩展至最高 131,072 tokens,需遵循官方配置并验证长文本质量。

文件说明

本次审核版本由 Hugging Face 官方 API 返回 15 个文件条目,总存储量约 42.33 GB。不同精度或量化格式的实际下载与运行占用可能明显不同。

上游文件元数据

  • .gitattributes1.53 KB
  • config.json728 B
  • generation_config.json239 B
  • LICENSE11.08 KB
  • merges.txt1.59 MB
  • model-00001-of-00005.safetensors3.72 GB
  • model-00002-of-00005.safetensors3.72 GB
  • model-00003-of-00005.safetensors3.69 GB
  • model-00004-of-00005.safetensors2.97 GB
  • model-00005-of-00005.safetensors1.16 GB
  • model.safetensors.index.json32.11 KB
  • README.md16.27 KB

硬件建议

建议以目标权重格式、上下文和并发量进行容量测试。显存或内存需求还受 KV Cache、推理框架、量化方法和是否使用 CPU/GPU 混合卸载影响,本站不以单一数字作保证。

注意事项

模型不能保证事实正确或工具调用安全。长上下文扩展、思考模式和高并发都会增加资源与等待时间;任何自动执行动作都应设置最小权限和人工确认点。

获取、校验与交付

咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。