中文简介
该数据集汇集多种模型生成的推理链,并筛选为约 5K token 以内的训练样本,面向小型语言模型的推理、代码和智能体训练。上游说明数据来自多个公开仓库并提供来源占比;使用前需要核对每个来源的数据许可、去重情况、推理痕迹质量和潜在污染。
上游模型卡 / 数据集卡
该数据集汇集多种模型生成的推理链,并筛选为约 5K token 以内的训练样本,面向小型语言模型的推理、代码和智能体训练。上游说明数据来自多个公开仓库并提供来源占比;使用前需要核对每个来源的数据许可、去重情况、推理痕迹质量和潜在污染。
Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs. The dataset has these columns for users to filter out: repo_id tok_len user thought_trace assistant ChatML Repositories Our reasoning corpus is a mix of carefully combining many smaller repositories. Our team broke down the amazing repositories we used to make this reasoning corpus: glaiveai/reasoning-v1-20m · 19.52% · 1,747,125,267 tokens · 1,053,837 rows PrimeIntellect/INTELLECT-3-SFT openreasoning_science · 9.87% · 883,229,296 tokens · 298,554 rows PrimeIntellect/INTELLECT-3-SFT am_chat · 8.28% · 741,154,467 tokens · 407,186 rows BAAI/OpenSeek-Synthetic-Reasoning-Data-Examples CC · 6.34% · 567,678,235 tokens · 673,768 rows nvidia/Nemotron-Cascade-SFT-Stage-1 general · 4.93% · 441,433,940 tokens · 305,200 rows open-thoughts/OpenThoughts2-1M · 4.78% · 428,267,556 tokens · 167,709 rows PrimeIntellect/SYNTHETIC-1-SFT-Data · 4.67% · 418,178,793 tokens · 207,147 rows Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B · 3.34% · 298,614,395 tokens · 181,183 rows Jackrong/LogicMind-Chat-Reasoning-SFT-300K · 3.12% · 278,997,572 tokens · 133,868 rows nvidia/Nemotron-Cascade-SFT-Stage-2 general · 2.81% · 251,628,095 tokens · 148,746 rows GeneralReasoning/GeneralThought-430K · 2.64% · 236,437,728 tokens · 133,755 rows Jackrong/GLM-5.1-Reasoning-1M-Cleaned main · 2.44% · 218,086,210 tokens · 76,152 rows allenai/Dolci-Think-SFT-7B · 2.41% · 216,086,528 tokens · 138,901 rows allenai/Dolci-Think-SFT-32B · 2.35% · 210,488,743 tokens · 135,400 rows nvidia/Nemotron-Cascade-SFT-Stage-1 math · 1.85% · 165,470,847 tokens · 55,090 rows ianncity/KIMI-K2.5-1000000x PHD-Science · 1.77% · 158,232,711 tokens · 45,433 rows ianncity/Hunter-Al
上游文件元数据
.gitattributes2.49 KBdataset.jsonl63.95 GBREADME.md55.51 KB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。