数据集

fineweb

HuggingFaceFW/fineweb

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:odc-by 上游版本:9bb295ddab0e

中文简介

FineWeb 是 Hugging Face 发布的大规模英文网页预训练语料,上游称总量约 15 万亿 token,并提供按 Common Crawl 批次拆分的数据及清洗、去重和评测说明。数据规模极大,使用时需关注网页来源权利、个人信息、偏差、过滤策略和 ODC-By 许可要求。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

FineWeb 是 Hugging Face 发布的大规模英文网页预训练语料,上游称总量约 15 万亿 token,并提供按 Common Crawl 批次拆分的数据及清洗、去重和评测说明。数据规模极大,使用时需关注网页来源权利、个人信息、偏差、过滤策略和 ODC-By 许可要求。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer Table of Contents 🍷 FineWeb What is it? What is being released? Changelog How to download and use 🍷 FineWeb + Using 🏭 `datatrove` + Using `huggingface_hub` + Using `datasets` Breakdown by dump/crawl Dataset performance evaluation and ablations + Hyper-parameters for ablation models + Ablation evaluation benchmarks + Comparison with other datasets Dataset card for 🍷 FineWeb Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information + Future work + Citation Information What is it? The 🍷 FineWeb dataset consists of more than **18.5T tokens** (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 `datatrove` library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a release of the **full dataset** under the **ODC-By 1.0 license**. However, by carefully adding additional filtering steps, we managed to push the performance of 🍷 FineWeb well above that of the original 🦅 RefinedWeb, and models trained on our dataset also outperform models trained on other commonly used high quality web datasets (like C4, Dolma-v1.6, The Pile, SlimPajama, RedPajam2) on our aggregate group of benchmark tasks. That said, we think there is still room for additional filtering and improvement and intend to continue exploring how to improve the dataset quality in coming versions of 🍷 FineWeb. What is being released? Along with the dataset, which includes all CommonCrawl dumps since 2013, we also share all the code needed to fully reproduce our p

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes2.25 KB
  • data/CC-MAIN-2013-20/000_00000.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00001.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00002.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00003.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00004.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00005.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00006.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00007.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00008.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00009.parquet2.00 GB
  • data/CC-MAIN-2013-20/000_00010.parquet2.00 GB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。