中文简介
FineWeb 是 Hugging Face 发布的大规模英文网页预训练语料,上游称总量约 15 万亿 token,并提供按 Common Crawl 批次拆分的数据及清洗、去重和评测说明。数据规模极大,使用时需关注网页来源权利、个人信息、偏差、过滤策略和 ODC-By 许可要求。
上游模型卡 / 数据集卡
FineWeb 是 Hugging Face 发布的大规模英文网页预训练语料,上游称总量约 15 万亿 token,并提供按 Common Crawl 批次拆分的数据及清洗、去重和评测说明。数据规模极大,使用时需关注网页来源权利、个人信息、偏差、过滤策略和 ODC-By 许可要求。
🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer Table of Contents 🍷 FineWeb What is it? What is being released? Changelog How to download and use 🍷 FineWeb + Using 🏭 `datatrove` + Using `huggingface_hub` + Using `datasets` Breakdown by dump/crawl Dataset performance evaluation and ablations + Hyper-parameters for ablation models + Ablation evaluation benchmarks + Comparison with other datasets Dataset card for 🍷 FineWeb Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information + Future work + Citation Information What is it? The 🍷 FineWeb dataset consists of more than **18.5T tokens** (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 `datatrove` library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a release of the **full dataset** under the **ODC-By 1.0 license**. However, by carefully adding additional filtering steps, we managed to push the performance of 🍷 FineWeb well above that of the original 🦅 RefinedWeb, and models trained on our dataset also outperform models trained on other commonly used high quality web datasets (like C4, Dolma-v1.6, The Pile, SlimPajama, RedPajam2) on our aggregate group of benchmark tasks. That said, we think there is still room for additional filtering and improvement and intend to continue exploring how to improve the dataset quality in coming versions of 🍷 FineWeb. What is being released? Along with the dataset, which includes all CommonCrawl dumps since 2013, we also share all the code needed to fully reproduce our p
上游文件元数据
.gitattributes2.25 KBdata/CC-MAIN-2013-20/000_00000.parquet2.00 GBdata/CC-MAIN-2013-20/000_00001.parquet2.00 GBdata/CC-MAIN-2013-20/000_00002.parquet2.00 GBdata/CC-MAIN-2013-20/000_00003.parquet2.00 GBdata/CC-MAIN-2013-20/000_00004.parquet2.00 GBdata/CC-MAIN-2013-20/000_00005.parquet2.00 GBdata/CC-MAIN-2013-20/000_00006.parquet2.00 GBdata/CC-MAIN-2013-20/000_00007.parquet2.00 GBdata/CC-MAIN-2013-20/000_00008.parquet2.00 GBdata/CC-MAIN-2013-20/000_00009.parquet2.00 GBdata/CC-MAIN-2013-20/000_00010.parquet2.00 GB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。