中文简介
allenai/molmo-motion-1m 是 allenai 在 Hugging Face 发布的数据集仓库,当前卡片主要关联机器人任务。本站固定记录上游版本 c3dec07d796d,同步到 82 个文件,文件元数据合计 265.94 GB;访问状态为“公开”,许可证字段为“mixed-per-subdataset”。
上游模型卡 / 数据集卡
【README 卡片中文译介】
以下内容依据上游仓库 allenai/molmo-motion-1m 的固定版本 README、卡片字段和文件元数据进行中文结构化整理。它保留便于选型的关键信息,完整技术细节、代码示例和许可文本请同时查看上游原文。
【资源概览】
allenai/molmo-motion-1m 是 allenai 在 Hugging Face 发布的数据集仓库,当前卡片主要关联机器人任务。本站固定记录上游版本 c3dec07d796d,同步到 82 个文件,文件元数据合计 265.94 GB;访问状态为“公开”,许可证字段为“mixed-per-subdataset”。
【适用方向】
适合评估具身智能、机器人学习、自动驾驶或强化学习研究流程。 正式使用前应先抽样检查字段、语言、重复项、敏感信息和训练/测试划分,避免只依据仓库名称判断数据质量。
【版本与规格】
固定版本:c3dec07d796d。 任务标签:机器人任务。 语言字段:en。 数据集卡名称:MolmoMotion-1M。 规模分类:1M<n<10M。 上游页面记录的最近更新时间:2026-06-18。
【文件信息】
上游 API 返回 82 个文件,文件元数据合计 265.94 GB。 已保存文件清单中的主要扩展名:.tar 37 个、.json 27 个、.md 8 个、.py 8 个、.txt 1 个、无扩展名 1 个。 当前较大的文件包括:xperience/tracks/tracks-0002.tar(10.08 GB);xperience/tracks/tracks-0001.tar(10.08 GB);xperience/tracks/tracks-0000.tar(10.08 GB)。
【运行或处理环境】
仓库文件元数据合计 265.94 GB。仅按下载、解压和临时文件估算,建议至少预留约 319.13 GB 可用磁盘;实际需求还取决于压缩率、缓存、数据转换和训练副本,正式处理前应以完整文件清单重新测算。
【使用边界】
许可证显示为“mixed-per-subdataset”,具体用途限制、附加政策和归属要求仍以上游原文为准。 数据集还需核对样本来源、隐私与个人信息、标签质量、地域偏差及下游模型的合规要求。
MolmoMotion-1M MolmoMotion-1M is a dataset of **3D point-trajectory annotations** curated across seven video corpora — ego-centric manipulation, real-world robot teleoperation, dynamic real-world scenes, and simulator renders. Each clip ships motion-filtered 3D tracks (and, for most datasets, 2D pixel tracks), a short action caption, per-frame camera, and a train/test split — all frame-aligned to the source video. Datasets We do not re-host the original videos, and for some datasets do not redistribute the upstream coordinates — we ship the annotations/tracks we produced, plus a `reconstruct_*.py` script to regenerate the rest from the original source. **Check each upstream dataset's license before downloading.** | Dataset | Domain | Upstream source | Ships | Reconstruct | |---|---|---|---|---| | egodex | ego-centric manip | apple/EgoDex | tracks + camera | videos | | ytvis | natural-video object tracks | YouTube-VIS 2021 | tracks + camera | videos | | hdepic | ego-centric kitchen | hd-epic.github.io | tracks + camera | videos | | xperience | ego-centric manip (gated) | ropedia-ai/xperience-10m | tracks | videos + camera | | stereo4d | dynamic VR180 scenes | stereo4d.github.io | track index | tracks + camera + videos | | droid | real-world robot teleop | droid-dataset.github.io | tracks + camera | videos | | molmospaces | sim pick-and-place | Ai2 internal sim | everything | — | Each dataset's `README.md` is **authoritative** for its schema and its exact reconstruction procedure. Download Unpack The per-clip directories (`tracks/`, `camera/`, and molmospaces `videos/` + `robot_trajectories/`) ship as ~10 GB tar shards inside those folders. Extract each dataset's shards from its folder: → each shard extracts into its own `tracks/` / `camera/` / `videos/` (stereo4d has no shards). The resulting layout is the tree below. Data format ` /annotations/` holds the JSONs — the **single source of truth**; don't enumerate the `tracks/` directory directly. ` _clips.json` — one e
适用场景
适合评估具身智能、机器人学习、自动驾驶或强化学习研究流程。 正式使用前应先抽样检查字段、语言、重复项、敏感信息和训练/测试划分,避免只依据仓库名称判断数据质量。
规模与格式
固定版本:c3dec07d796d。 任务标签:机器人任务。 语言字段:en。 数据集卡名称:MolmoMotion-1M。 规模分类:1M<n<10M。 上游页面记录的最近更新时间:2026-06-18。
文件说明
上游 API 返回 82 个文件,文件元数据合计 265.94 GB。 已保存文件清单中的主要扩展名:.tar 37 个、.json 27 个、.md 8 个、.py 8 个、.txt 1 个、无扩展名 1 个。 当前较大的文件包括:xperience/tracks/tracks-0002.tar(10.08 GB);xperience/tracks/tracks-0001.tar(10.08 GB);xperience/tracks/tracks-0000.tar(10.08 GB)。
上游文件元数据
.gitattributes3.49 KBdroid/annotations/droid_clips.json6.68 MBdroid/annotations/droid_split.json6.68 MBdroid/annotations/droid_videos_index.json2.94 MBdroid/camera/camera-0000.tar45.91 MBdroid/README.md4.37 KBdroid/reconstruct_videos.py6.27 KBdroid/tracks/tracks-0000.tar2.12 GBegodex/annotations/egodex_clips.json21.04 MBegodex/annotations/egodex_hand_split.json43.20 MBegodex/annotations/egodex_paired_index.json3.40 MBegodex/annotations/egodex_split.json21.04 MB
硬件建议
仓库文件元数据合计 265.94 GB。仅按下载、解压和临时文件估算,建议至少预留约 319.13 GB 可用磁盘;实际需求还取决于压缩率、缓存、数据转换和训练副本,正式处理前应以完整文件清单重新测算。
注意事项
许可证显示为“mixed-per-subdataset”,具体用途限制、附加政策和归属要求仍以上游原文为准。 数据集还需核对样本来源、隐私与个人信息、标签质量、地域偏差及下游模型的合规要求。
获取、校验与交付
咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。
本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。