模型

Qwen3 1.7B

Qwen/Qwen3-1.7B

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:已人工审核 许可证:Apache License 2.0 上游版本:70d244cc86cc

中文简介

Qwen3 1.7B 在轻量部署与基础生成能力之间提供更宽裕的空间,支持思考和非思考模式。它适合希望控制本地资源成本,同时比 0.6B 档获得更完整表达能力的研发与教学场景。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Qwen3 1.7B 是 Qwen3 系列的轻量稠密模型,官方说明其同时支持思考模式和非思考模式。

本版本原生上下文长度为 32,768 tokens。与 0.6B 档相比,它为本地原型和基础生成任务提供更大的参数容量,但仍需根据具体任务验证效果。

部署结论不能仅按参数量推断。应结合精度或量化格式、上下文长度、并发量和推理框架评估真实资源需求,并以上游许可证和固定版本说明为准。

译文已人工复核 · 本站基于上游 README 整理并人工复核 · 2026-07-23 23:47

Qwen3-1.7B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: **Uniquely support of seamless switching between thinking mode** (for complex logical reasoning, math, and coding) and **non-thinking mode** (for efficient, general-purpose dialogue) **within single model**, ensuring optimal performance across various scenarios. **Significantly enhancement in its reasoning capabilities**, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning. **Superior human preference alignment**, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience. **Expertise in agent capabilities**, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks. **Support of 100+ languages and dialects** with strong capabilities for **multilingual instruction following** and **translation**. Model Overview Qwen3-1.7B** has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 1.7B Number of Paramaters (Non-Embedding): 1.4B Number of Layers: 28 Number of Attention Heads (GQA): 16 for Q and 8 for KV Context Length: 32,768 For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. [!TIP] If you encounter significant endless repetitions, please refer to the Best Practices section for optimal sampling parameters, and s

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合本地助手原型、基础信息抽取、短文本生成、教学实验和边缘侧可行性验证。用于专业知识或面向客户的内容前,应建立任务级评测和人工复核。

模型参数

约 1.7B 参数的稠密模型,原生上下文长度为 32,768 tokens。支持思考与非思考模式,可根据质量、速度和 Token 成本选择。

文件说明

本次固定版本由 Hugging Face 官方 API 返回 12 个文件条目,总存储量约 3.80 GB。文件元数据仅用于版本和交付核对,不是运行资源承诺。

上游文件元数据

  • .gitattributes1.53 KB
  • config.json726 B
  • generation_config.json239 B
  • LICENSE11.08 KB
  • merges.txt1.59 MB
  • model-00001-of-00002.safetensors3.20 GB
  • model-00002-of-00002.safetensors593.50 MB
  • model.safetensors.index.json25.00 KB
  • README.md13.64 KB
  • tokenizer.json10.89 MB
  • tokenizer_config.json9.50 KB
  • vocab.json2.65 MB

硬件建议

通常比中大型模型更容易在本地环境验证,但精度、量化、上下文、并发和框架会显著改变内存与速度。应先以目标格式和真实提示词进行基准测试。

注意事项

小参数模型在复杂推理、长链条指令和专业知识场景中可能不稳定。不要把单次示例效果当作生产结论,也不要在未评估前输入敏感或受限制数据。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。