中文简介
这是一个由多个来源汇总的超大规模蒸馏数据集合,上游宣称包含 1800 万以上样本、数十个来源及多个类别。仓库名称和数据卡含有大量第三方模型品牌与规模声明,本站未验证其真实来源、授权链和数据质量;使用前应逐项审查来源、许可、去重和合规风险。
上游模型卡 / 数据集卡
这是一个由多个来源汇总的超大规模蒸馏数据集合,上游宣称包含 1800 万以上样本、数十个来源及多个类别。仓库名称和数据卡含有大量第三方模型品牌与规模声明,本站未验证其真实来源、授权链和数据质量;使用前应逐项审查来源、许可、去重和合规风险。
📖 The Open Distillation Codex 🌌 *The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense* 🌌 Where 73 open-source minds converge into one unified stream of intelligence** `18M+ Distilled Signals` · `7,090 Raw GitHub Repositories` · `8 Curated Categories` · `~76 GB+` *"We did not write this dataset. We assembled it.* *Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing.* *Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."* 📌 Table of Contents | # | Section | Description | |---|---|---| | 1 | 📊 Dataset Summary | High-level overview & value proposition | | 2 | 🗂️ Directory Structure | ASCII tree + folder explanation | | 3 | 🌐 Data Sources | All 73 sources with full attribution | | 4 | 🛡️ Cybersecurity Deep Dive: Attack & Defense | Importance, attack traces, defense, exploit analysis | | 5 | 🛠️ How to Use & Train | Loading, streaming, training scripts | | 6 | 🔐 Licensing & Limitations | License, intended use, limitations | | 7 | 📜 Changelog | Version history | 📊 Dataset Summary 🎯 The Numbers That Matter | Metric | Value | Status | |:---:|:---:|:---:| | **Total Storage** | `76 GB+` | ✅ Verified | | **JSONL Data Shards** | `516` | ✅ Verified | | **Archive Files (tar.gz)** | `7,090` | ✅ Verified | | **Source Datasets** | `73` | ✅ Verified | | **Categories** | `8` | ✅ Verified | | **Total Samples** | `18M+` | ✅ Verified | | **Largest Source** | `8.15M` (Vibe-Coding-Instruct-V2) | ✅ | | **Archive Size** | `~64 GB` (compressed GitHub repos) | ✅ | | **Cybersecurity Sources** | `6` | ✅ | | **Cybersecurity Data Size** | `~2.6 GB` | ✅ | 🌟 Why "Ultimate Distilled"? This dataset is not a raw scrape. Every sample has been **distilled through a unified extraction pipeline**: 💎 Value to the Open-Source AI Community | 🎯 For... | 📦 This dataset provides... | |---|---| | **Model Trainers** | Single `load_dataset()`
上游文件元数据
.gitattributes1.68 KBarchives/0-chi__sonaure-lp.tar.gz49.63 KBarchives/00MB__bitcoin_trading_bot.tar.gz249.99 KBarchives/0101-agents__plugins.tar.gz16.15 KBarchives/01LETO__Avorex.tar.gz2.09 MBarchives/0Do7__ascii-games.tar.gz25.37 KBarchives/0scarito__0scarito.tar.gz25.17 KBarchives/0x-CryptoPriest__scar.tar.gz96.20 KBarchives/0x000x7f__0x000x7f.tar.gz13.12 KBarchives/0x000x7f__codex-cli-mcp-bridge.tar.gz213.44 KBarchives/0x101__lakewatch.tar.gz855.65 KBarchives/0x8801__neticle.tar.gz479.88 KB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。