中文简介
ThinkingCap-Qwen3.6-27B 是基于 Qwen3.6-27B 的微调模型,目标是在尽量保持回答质量与风格的同时减少推理 token。上游称其在多类推理、对话、系统指令、安全、数学、代码和智能体测试中进行了多种随机种子评估;节省比例和质量保持程度需在真实任务上复测。
上游模型卡 / 数据集卡
ThinkingCap-Qwen3.6-27B 是基于 Qwen3.6-27B 的微调模型,目标是在尽量保持回答质量与风格的同时减少推理 token。上游称其在多类推理、对话、系统指令、安全、数学、代码和智能体测试中进行了多种随机种子评估;节省比例和质量保持程度需在真实任务上复测。
ThinkingCap: Qwen 3.6 27B Capability of Qwen3.6-27B with **50% less** thinking tokens on average, and over **90% less** in best cases. Achieved via finetuning Qwen3.6-27B (Qwen Team, 2026) with state-of-the-art algorithms on a curated set of problems of various domains and difficulty. We designed the finetuning to be as minimally invasive as possible, preserving all of the original answer quality and style of Qwen, while being more token efficient. Check the blogpost for more details. We rigorously evaluate the resulting checkpoint across general reasoning, non-reasoning multiple-choice question answering, everyday multi-turn conversations, system prompt adherence, safety, math, code and agentic use cases. Due to the high variability of reasoning quality at Qwen-recommended sampling temperature 1.0, we run each benchmark with multiple seeds and do statistical significance testing on all the results. We evaluate both in domain (holdout parts of selected datasets included in training) and out of domain. Out-of-domain token efficiency Benchmark Accuracy Thinking tokens Base Ours Base Ours Reduction Knowledge & reasoning GPQA-Diamond 85.5 ±1.4 83.8 ±1.9 10,777 3,351 ↓ 67.8% SuperGPQA 64.0 ±0.2 64.0 ±0.1 8,246 3,384 ↓ 58.4% MMLU-Pro 85.9 ±0.2 85.4 ±0.2 3,455 1,290 ↓ 53.7% MMLU-Redux 93.9 ±0.1 93.9 ±0.1 947 406 ↓ 44.8% C-Eval 90.6 ±0.7 90.3 ±0.6 1,279 663 ↓ 47.1% Math & code HMMT (Nov 2025) 88.0 ±3.7 84.7 ±3.7 39,277 27,388 ↓ 38.0% LiveCodeBench 80.7 ±0.6 84.3 ±1.0 15,744 10,158 ↓ 41.1% Long-context & multimodal LongBench v2 62.6 ±3.6 60.2 ±1.7 1,765 1,091 ↓ 39.1% RealWorldQA 82.4 ±0.7 81.9 ±1.2 2,959 913 ↓ 48.5% AA-LCR 76.2 ±3.0 74.2 ±2.2 2,455 1,337 ↓ 45.5% Instruction following & agentic System-prompt adherence 80.6 ±1.2 81.5 ±1.8 1,737 976 ↓ 40.0% Claw-Eval think/task 87.0 ±1.9 84.4 ±1.2 919 689 ↓ 25.2% Macro average 81.5 80.7 — — ↓ 45.8% Claw-Eval thinking tokens are per-task (agentic; not a single-turn trace). Settings** **Models:** base `Qwen/Qwen3.
上游文件元数据
.eval_results/gpqa.yaml169 B.eval_results/gsm8k.yaml164 B.eval_results/mmlu-pro.yaml173 B.gitattributes1.69 KBbc_capybara.png2.85 MBbottlecap-logo.png33.96 KBbudget_aggregate.png87.92 KBcap_header.png1.01 MBchat_template.jinja7.58 KBconfig.json3.60 KBconfiguration.json51 Bgeneration_config.json213 B
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。