数据集

fineweb-edu

HuggingFaceFW/fineweb-edu

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:odc-by 上游版本:87f09149ef47

中文简介

FineWeb-Edu 是从 FineWeb 中筛选出的英文教育网页语料,上游介绍的主要版本约有 1.3 万亿 token。筛选器使用模型生成的教育质量标注训练,用于保留更具教育价值的页面;使用时仍需关注网页来源权利、个人信息、偏差、过滤误差和 ODC-By 许可要求。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

FineWeb-Edu 是从 FineWeb 中筛选出的英文教育网页语料,上游介绍的主要版本约有 1.3 万亿 token。筛选器使用模型生成的教育质量标注训练,用于保留更具教育价值的页面;使用时仍需关注网页来源权利、个人信息、偏差、过滤误差和 ODC-By 许可要求。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper:** https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of **1.3T tokens** and **5.4T tokens** (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data. The Dataset Curation section details the process for creating the dataset. You can find a deduplicated version of FineWeb-edu in SmolLM-Corpus. We find that the deduplication of this dataset doesn't have any impact on model performance in our ablation setup (1.8B trained on 350B tokens). What is being released? Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification Changelog _Previous versions remain available in the branch `version name`._ **v1.4.0 (11-07-2025):** Added 6 new snapshots: `CC-MAIN-2025-05`, `CC-MAIN-2025-08`, `CC-MAIN-2025-13`, `CC-MAIN-2025-18`, `CC-MAIN-2025-21`, and `CC-MAIN-2025-26` (January to June 2025) **v1.3.0 (31-01-2025):** Fixed an issue with some dumps where some documents hadn't been processed: `CC-MAIN-2024-10`, `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2024-42`, `CC-MAIN-2024-46` -- they now contain more data (~35B additional tokens). **v1.2.0 (03-01-2025):** Added 9 new snapshots: `CC-MAIN-2024-18`, `CC-MAIN-2024-22`, `CC-MAIN-2024-26`, `CC-MAIN-2024-30`, `CC-MAIN-2024-33`, `CC-MAIN-2024-38`, `CC-MAIN-2

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes2.25 KB
  • data/CC-MAIN-2013-20/train-00000-of-00014.parquet2.21 GB
  • data/CC-MAIN-2013-20/train-00001-of-00014.parquet2.21 GB
  • data/CC-MAIN-2013-20/train-00002-of-00014.parquet2.20 GB
  • data/CC-MAIN-2013-20/train-00003-of-00014.parquet2.19 GB
  • data/CC-MAIN-2013-20/train-00004-of-00014.parquet2.19 GB
  • data/CC-MAIN-2013-20/train-00005-of-00014.parquet2.18 GB
  • data/CC-MAIN-2013-20/train-00006-of-00014.parquet2.18 GB
  • data/CC-MAIN-2013-20/train-00007-of-00014.parquet2.16 GB
  • data/CC-MAIN-2013-20/train-00008-of-00014.parquet2.15 GB
  • data/CC-MAIN-2013-20/train-00009-of-00014.parquet2.15 GB
  • data/CC-MAIN-2013-20/train-00010-of-00014.parquet2.14 GB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。