中文简介
TinyStories 是由 GPT-3.5 和 GPT-4 合成的英文短故事数据集,故事使用较小词汇表,常用于研究小型语言模型的生成与语言学习能力。仓库包含训练、验证及扩展版本资料;合成偏差、内容质量和 CDLA-Sharing 许可条件需要在使用前核对。
上游模型卡 / 数据集卡
TinyStories 是由 GPT-3.5 和 GPT-4 合成的英文短故事数据集,故事使用较小词汇表,常用于研究小型语言模型的生成与语言学习能力。仓库包含训练、验证及扩展版本资料;合成偏差、内容质量和 CDLA-Sharing 许可条件需要在使用前核对。
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of the stories together with metadata and the prompt that was used to create each story. TinyStoriesV2-GPT4-train.txt - Is a new version of the dataset that is based on generations by GPT-4 only (the original dataset also has generations by GPT-3.5 which are of lesser quality). It contains all the examples in TinyStories.txt which were GPT-4 generated as a subset (but is significantly larger). Evaluation_prompts.yaml: List of prompts used to evaluate our models (see paper)
上游文件元数据
.gitattributes2.69 KBdata/train-00000-of-00004-2d5a1467fff1081b.parquet237.21 MBdata/train-00001-of-00004-5852b56a2bd28fd9.parquet236.68 MBdata/train-00002-of-00004-a26307300439e943.parquet234.50 MBdata/train-00003-of-00004-d243063613e5a057.parquet236.50 MBdata/validation-00000-of-00001-869c898b519ad725.parquet9.53 MBEvaluation prompts.yaml11.48 KBMLP_input_output.npy781.25 MBREADME.md1.04 KBTinyStories-train.txt1.79 GBTinyStories-valid.txt18.55 MBTinyStories_all_data.tar.gz1.50 GB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。