模型

Cosmos3-Edge

nvidia/Cosmos3-Edge

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:other 上游版本:6f58f6b4c912

中文简介

Cosmos3-Edge 属于 NVIDIA Cosmos 3 物理 AI 世界模型系列,面向机器人、自动驾驶、工业和智能空间等场景。上游将 Cosmos 3 描述为可结合文本、图像、视频和动作轨迹生成视频、图像、音频或动作命令的全模态基础组件;具体许可和商业使用条件需查阅 NVIDIA 原文。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Cosmos3-Edge 属于 NVIDIA Cosmos 3 物理 AI 世界模型系列,面向机器人、自动驾驶、工业和智能空间等场景。上游将 Cosmos 3 描述为可结合文本、图像、视频和动作轨迹生成视频、图像、音频或动作命令的全模态基础组件;具体许可和商业使用条件需查阅 NVIDIA 原文。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

**Cosmos 3: Omnimodal World Models for Physical AI** Model Collection** | **Code** | **White Paper** | **Website** NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications. Model Overview: Cosmos3-Edge Description Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning. This model is ready for commercial and non-commercial use. Model Developer:** NVIDIA Model Versions Released on: 07/20/2026** Cosmos3-Edge: Given multimodal inputs including text, images, video, and action trajectories, generate coherent text, images, video, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications. Cosmos3-Edge-Policy-DROID: Given language instructions and visual observations from the DROID robot platform, generate robot action trajectories for manipulation and control tasks. Cosmos3-Super-Image2Video-4Step: Given one or more input images and optional text instructions, generate temporally coherent video sequences that are consistent with the provided visual content. Distilled from Cosmos3-Super-Image2Video using Improved Distribution Matching Distillation (DMD2), enabling high-quality generation in 4 steps. Cosmos3-Super-Text2Image-4Step: Given text input, generate high-fidelity images that are consistent with the provided description. Distilled from Cosmos3-Super-Text2Image using Improved Distribution Matching Distillation (DM

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes2.75 KB
  • assets/benchmark-image2video.png56.34 KB
  • assets/benchmark-overall.png110.32 KB
  • assets/edge_action_fd_umi_2chunk_output.mp432.11 KB
  • assets/edge_action_id_av_0_output.json15.38 KB
  • assets/edge_action_id_av_0_output.png197.55 KB
  • assets/edge_action_id_av_1_output.json15.21 KB
  • assets/edge_action_id_av_1_output.png154.04 KB
  • assets/edge_i2v_output.mp422.27 MB
  • assets/example_action_fd_umi_2chunk_output.mp435.21 KB
  • assets/example_action_fd_umi_action_chunks.json10.11 KB
  • assets/example_action_fd_umi_first_frame.png65.46 KB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。