数据集

SenseNova-Vision-Corpus-50M

sensenova/SenseNova-Vision-Corpus-50M

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:cc-by-nc-4.0 上游版本:4f144b7cae11

中文简介

SenseNova-Vision-Corpus-50M 是面向统一多模态生成的图文语料集合,上游同时提供英文与简体中文数据卡入口。许可证标记为 CC BY-NC 4.0,意味着默认限制商业使用;数据规模、图像来源、过滤方式和敏感内容处理必须在实际采用前进一步审核。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

SenseNova-Vision-Corpus-50M 是面向统一多模态生成的图文语料集合,上游同时提供英文与简体中文数据卡入口。许可证标记为 CC BY-NC 4.0,意味着默认限制商业使用;数据规模、图像来源、过滤方式和敏感内容处理必须在实际采用前进一步审核。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Vision as Unified Multimodal Generation English | 简体中文 This repository contains the dataset for the paper Vision as Unified Multimodal Generation. SenseNova Vision Corpus 50M Overview SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing public vision datasets: annotations are often incomplete, inconsistent across tasks, or not directly decodable into training-ready multimodal supervision. SN-VC-50M organizes open-source visual data into four task families: **structured visual understanding**, **segmentation**, **dense geometric prediction**, and **multi-view visual geometry**. Across these families, the release contains **73 dataset-task entries** covering **10 task types**. The released corpus includes: **18.9M frames** for structured visual understanding **1.3M frames** for segmentation **17.3M frames** for dense geometric prediction **12.5M frames** for multi-view visual geometry To construct training-compatible supervision, we use task-specialized curation pipelines. For structured understanding, we adapt the Rex-Omni data construction pipeline for detection- and OCR-style sample generation. For dense geometry, we leverage MoGe-2 to densify sparse depth and surface-normal annotations and improve scene diversity. For multi-view scenarios, we use LingBot-Depth to complement incomplete sparse depth information. For segmentation, we apply strict alignment checks to ensure consistency between textual region descriptions, color legends, and segmentation masks. To avoid redistributing duplicated raw RGB images from public source datasets, the JSONL training examples retain the corresponding **relative file paths** instead of duplicating all original RGB assets. Users need to align the local dataset root directory with the image file paths recorded in the corresponding JSONL files to

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes7.85 KB
  • Dataset_download_description.md4.75 KB
  • dense_geometric_prediction/coco_total_unified_depth_edit_regenerated.jsonl30.91 MB
  • dense_geometric_prediction/coco_total_unified_normal_edit_regenerated.jsonl31.92 MB
  • dense_geometric_prediction/object365_total_unified_depth_edit_regenerated.jsonl546.63 MB
  • dense_geometric_prediction/object365_total_unified_normal_edit_regenerated.jsonl557.26 MB
  • dense_geometric_prediction/sa_1b_total_unified_depth_edit_regenerated.jsonl1.70 GB
  • dense_geometric_prediction/sa_1b_total_unified_normal_edit_regenerated.jsonl1.75 GB
  • dense_geometric_prediction/scannetpp_total_unified_depth_edit_regenerated.jsonl331.96 MB
  • dense_geometric_prediction/scannetpp_total_unified_normal_edit_regenerated.jsonl338.55 MB
  • dense_geometric_prediction/taskonomy_total_unified_depth_edit_regenerated.jsonl853.85 MB
  • dense_geometric_prediction/taskonomy_total_unified_normal_edit_regenerated.jsonl870.02 MB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。