中文简介
SenseNova-Vision-Corpus-50M 是面向统一多模态生成的图文语料集合,上游同时提供英文与简体中文数据卡入口。许可证标记为 CC BY-NC 4.0,意味着默认限制商业使用;数据规模、图像来源、过滤方式和敏感内容处理必须在实际采用前进一步审核。
上游模型卡 / 数据集卡
SenseNova-Vision-Corpus-50M 是面向统一多模态生成的图文语料集合,上游同时提供英文与简体中文数据卡入口。许可证标记为 CC BY-NC 4.0,意味着默认限制商业使用;数据规模、图像来源、过滤方式和敏感内容处理必须在实际采用前进一步审核。
Vision as Unified Multimodal Generation English | 简体中文 This repository contains the dataset for the paper Vision as Unified Multimodal Generation. SenseNova Vision Corpus 50M Overview SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing public vision datasets: annotations are often incomplete, inconsistent across tasks, or not directly decodable into training-ready multimodal supervision. SN-VC-50M organizes open-source visual data into four task families: **structured visual understanding**, **segmentation**, **dense geometric prediction**, and **multi-view visual geometry**. Across these families, the release contains **73 dataset-task entries** covering **10 task types**. The released corpus includes: **18.9M frames** for structured visual understanding **1.3M frames** for segmentation **17.3M frames** for dense geometric prediction **12.5M frames** for multi-view visual geometry To construct training-compatible supervision, we use task-specialized curation pipelines. For structured understanding, we adapt the Rex-Omni data construction pipeline for detection- and OCR-style sample generation. For dense geometry, we leverage MoGe-2 to densify sparse depth and surface-normal annotations and improve scene diversity. For multi-view scenarios, we use LingBot-Depth to complement incomplete sparse depth information. For segmentation, we apply strict alignment checks to ensure consistency between textual region descriptions, color legends, and segmentation masks. To avoid redistributing duplicated raw RGB images from public source datasets, the JSONL training examples retain the corresponding **relative file paths** instead of duplicating all original RGB assets. Users need to align the local dataset root directory with the image file paths recorded in the corresponding JSONL files to
上游文件元数据
.gitattributes7.85 KBDataset_download_description.md4.75 KBdense_geometric_prediction/coco_total_unified_depth_edit_regenerated.jsonl30.91 MBdense_geometric_prediction/coco_total_unified_normal_edit_regenerated.jsonl31.92 MBdense_geometric_prediction/object365_total_unified_depth_edit_regenerated.jsonl546.63 MBdense_geometric_prediction/object365_total_unified_normal_edit_regenerated.jsonl557.26 MBdense_geometric_prediction/sa_1b_total_unified_depth_edit_regenerated.jsonl1.70 GBdense_geometric_prediction/sa_1b_total_unified_normal_edit_regenerated.jsonl1.75 GBdense_geometric_prediction/scannetpp_total_unified_depth_edit_regenerated.jsonl331.96 MBdense_geometric_prediction/scannetpp_total_unified_normal_edit_regenerated.jsonl338.55 MBdense_geometric_prediction/taskonomy_total_unified_depth_edit_regenerated.jsonl853.85 MBdense_geometric_prediction/taskonomy_total_unified_normal_edit_regenerated.jsonl870.02 MB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。