中文简介
Qwen3 1.7B 在轻量部署与基础生成能力之间提供更宽裕的空间,支持思考和非思考模式。它适合希望控制本地资源成本,同时比 0.6B 档获得更完整表达能力的研发与教学场景。
上游模型卡 / 数据集卡
Qwen3 1.7B 是 Qwen3 系列的轻量稠密模型,官方说明其同时支持思考模式和非思考模式。
本版本原生上下文长度为 32,768 tokens。与 0.6B 档相比,它为本地原型和基础生成任务提供更大的参数容量,但仍需根据具体任务验证效果。
部署结论不能仅按参数量推断。应结合精度或量化格式、上下文长度、并发量和推理框架评估真实资源需求,并以上游许可证和固定版本说明为准。
Qwen3-1.7B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-1.7B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 1.7B Number of Paramaters (Non-Embedding): 1.4B Number of Layers: 28 Number of Attention Heads (GQA): 16 for Q and 8 for KV Context Length: 32,768 For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. [!TIP] If you encounter significant endless repetitions, please refer to the Best Practices section for optimal sampling parameters, and s
适用场景
适合本地助手原型、基础信息抽取、短文本生成、教学实验和边缘侧可行性验证。用于专业知识或面向客户的内容前,应建立任务级评测和人工复核。
模型参数
约 1.7B 参数的稠密模型,原生上下文长度为 32,768 tokens。支持思考与非思考模式,可根据质量、速度和 Token 成本选择。
文件说明
本次固定版本由 Hugging Face 官方 API 返回 12 个文件条目,总存储量约 3.80 GB。文件元数据仅用于版本和交付核对,不是运行资源承诺。
上游文件元数据
.gitattributes1.53 KBconfig.json726 Bgeneration_config.json239 BLICENSE11.08 KBmerges.txt1.59 MBmodel-00001-of-00002.safetensors3.20 GBmodel-00002-of-00002.safetensors593.50 MBmodel.safetensors.index.json25.00 KBREADME.md13.64 KBtokenizer.json10.89 MBtokenizer_config.json9.50 KBvocab.json2.65 MB
硬件建议
通常比中大型模型更容易在本地环境验证,但精度、量化、上下文、并发和框架会显著改变内存与速度。应先以目标格式和真实提示词进行基准测试。
注意事项
小参数模型在复杂推理、长链条指令和专业知识场景中可能不稳定。不要把单次示例效果当作生产结论,也不要在未评估前输入敏感或受限制数据。
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。