中文简介
FineWeb 是 Hugging Face 团队发布的大规模英文网页预训练语料,固定数据集卡说明包含超过 18.5T token,来源覆盖多个 Common Crawl 抓取批次。它适合大规模语言模型预训练研究,但不是“已完全清除个人信息和有害内容”的可直接商用数据包。
上游模型卡 / 数据集卡
数据定位:FineWeb 是大规模英文网页预训练语料,固定卡片给出 18.5T+ token,并按 Common Crawl 批次组织。它提供多个较小样本,不需要所有项目都获取全量。
选择建议:先明确训练 token 预算、年份范围和过滤目标,再选择 10B、100B、350B 样本或具体 crawl。API used_storage 约 106.77 TB,不应直接作为某个子集的交付大小。
合规边界:数据卡明确提示仍可能含个人信息、有害内容和偏差。ODC-By 归属、Common Crawl 条款、删除请求与二次过滤都应纳入项目治理;本站可协助固定配置和文件清单,不承诺数据天然适合商用训练。
🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer Table of Contents 🍷 FineWeb What is it? What is being released? Changelog How to download and use 🍷 FineWeb + Using 🏭 `datatrove` + Using `huggingface_hub` + Using `datasets` Breakdown by dump/crawl Dataset performance evaluation and ablations + Hyper-parameters for ablation models + Ablation evaluation benchmarks + Comparison with other datasets Dataset card for 🍷 FineWeb Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information + Future work + Citation Information What is it? The 🍷 FineWeb dataset consists of more than **18.5T tokens** (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 `datatrove` library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a release of the **full dataset** under the **ODC-By 1.0 license**. However, by carefully adding additional filtering steps, we managed to push the performance of 🍷 FineWeb well above that of the original 🦅 RefinedWeb, and models trained on our dataset also outperform models trained on other commonly used high quality web datasets (like C4, Dolma-v1.6, The Pile, SlimPajama, RedPajam2) on our aggregate group of benchmark tasks. That said, we think there is still room for additional filtering and improvement and intend to continue exploring how to improve the dataset quality in coming versions of 🍷 FineWeb. What is being released? Along with the dataset, which includes all CommonCrawl dumps since 2013, we also share all the code needed to fully reproduce our p
适用场景
适合英文语言模型预训练、数据过滤与去重研究、不同网页抓取批次的消融实验。多数团队应先使用 10B、100B 或 350B token 样本,或按单个 crawl 流式读取;只有明确训练规模、数据治理和存储预算后才考虑全量。
规模与格式
固定版本:9bb295ddab0e05d785b879661af7260fed5140fc。数据集卡:英文、超过 18.5T token、Parquet,按 Common Crawl dump 组织,并提供约 10B、100B、350B token 的样本配置。本站记录的规模分类字段为 n>1T。
文件说明
Hugging Face API 记录 28,147 个文件,used_storage 字段约 106.77 TB;本站文件元数据清单已截断。该存储字段与数据集卡按 crawl 或 sample 给出的下载体量不是同一口径。需求确认必须明确配置、crawl、样本规模和 Parquet 分片,不能只写“下载 FineWeb 全部”。
上游文件元数据
.gitattributes2.25 KBdata/CC-MAIN-2013-20/000_00000.parquet2.00 GBdata/CC-MAIN-2013-20/000_00001.parquet2.00 GBdata/CC-MAIN-2013-20/000_00002.parquet2.00 GBdata/CC-MAIN-2013-20/000_00003.parquet2.00 GBdata/CC-MAIN-2013-20/000_00004.parquet2.00 GBdata/CC-MAIN-2013-20/000_00005.parquet2.00 GBdata/CC-MAIN-2013-20/000_00006.parquet2.00 GBdata/CC-MAIN-2013-20/000_00007.parquet2.00 GBdata/CC-MAIN-2013-20/000_00008.parquet2.00 GBdata/CC-MAIN-2013-20/000_00009.parquet2.00 GBdata/CC-MAIN-2013-20/000_00010.parquet2.00 GB
硬件建议
建议优先采用流式或按 crawl 读取,并为 Parquet 扫描、去重、二次过滤和中间产物准备独立对象存储。全量处理通常需要分布式计算和远高于最终样本大小的临时空间。小团队从 10B token 样本开始更容易验证数据管线和合规流程。
注意事项
固定数据集卡明确说明仍很可能包含个人可识别信息,也可能残留有害、偏见或版权敏感内容。ODC-By 要求归属,使用还受 Common Crawl 条款约束;训练前应建立删除请求、来源追踪、敏感信息扫描和用途审查机制。
获取、校验与交付
咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。