数据集

PerceptionBench

moonshotai/PerceptionBench

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:AI 辅助整理并检查 许可证:Creative Commons Attribution-NonCommercial 4.0 International 上游版本:6ba8c3135c76

中文简介

PerceptionBench 是 Moonshot AI 发布的多模态模型原子视觉感知评测集,包含 3,000 个短答案问题,覆盖视觉关系、计数、属性、深度、定位、比较、细粒度识别、上下文、OCR 与感知幻觉十类能力。它适合做诊断评测,不是通用视觉预训练语料。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

数据定位:PerceptionBench 是用于诊断多模态大模型原子视觉感知能力的评测集,共 3,000 条短答案问题和十类能力。它强调把感知错误与推理或知识错误分开,不应当作通用训练数据集描述。

文件与使用:固定仓库的核心 JSONL 约 1.52 GiB,图像内容内嵌在样本中。使用前应检查字段、图像解码、评分提示和目标模型输入格式,并从小样本开始验证评测链路。

许可与交付:许可证字段为 CC BY-NC 4.0,要求署名并限制商业用途,第三方图像与来源还需单独核对。橙子AI科技可协助固定版本、文件、README、许可证和校验信息确认,并按规模提供网盘或硬盘交付咨询。

已有简体中文译文 · Codex 基于固定版本上游 README 编写;待人工复核 · 2026-08-07 23:24

PerceptionBench PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models** Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the response of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved—no model reaches 60% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs. Dataset Statistics The released benchmark comprises **3,000 verified questions** across the **ten atomic perceptual capabilities**. By construction source, **1,800 (60%)** are atomic sub-questions decomposed from attributed failures on the source benchmarks, while the remaining **1,200 (40%)** are newly authored on supplemented images. The 3,000 questions are subsampled with capability-level balancing and difficulty stratification from the constructed portion of an in-house pool of **17,000+ verified samples**. Leaderboard Sixteen frontier MLLMs (ten proprietary,

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合比较多模态模型在单一感知能力上的短板、复现统一提示下的开放式短答案评测,并为 OCR、计数或定位等能力建立回归测试。使用前应先检查 JSONL 字段、图像编码、评分器与目标模型输入格式,再决定是否全量运行。

规模与格式

固定版本:6ba8c3135c7675ad6a5c141536a86b9460c70960。数据集卡:3,000 个已核验问题、十类原子视觉感知能力;1,800 条为既有基准失败案例的原子子问题,1,200 条为补充图像上的新编问题。任务字段为 visual-question-answering,规模分类为 1K<n<10K。

文件说明

固定 SHA 下共有 13 个文件,核心 PerceptionBench.jsonl 为 1,626,804,092 字节(约 1.52 GiB),另有 README 与说明图片;Hugging Face used_storage 约 3.03 GB。交付应以固定 SHA 的当前文件清单为准,并核对 JSONL 哈希与可解析性。

上游文件元数据

  • .gitattributes100 B
  • .gitignore18 B
  • images/arxiv_small.svg874 B
  • images/code_badge.svg1.67 KB
  • images/dataset_badge.svg677 B
  • images/fig_distribution.png696.17 KB
  • images/fig_teaser.png1.99 MB
  • images/github_small.svg1.52 KB
  • images/homepage_badge.svg648 B
  • images/kimi_small.png229.35 KB
  • images/paper_badge.svg1.24 KB
  • PerceptionBench.jsonl1.52 GB

硬件建议

数据文件本身不要求高端 GPU,但 JSONL 内嵌图像会增加解析、解码和缓存开销;全量多模型评测还需按 3,000 条样本乘以模型数量规划推理成本。建议先抽样验证评分器和字段,再分批运行并保留错误样本日志。

注意事项

仓库卡片许可证为 CC BY-NC 4.0,要求署名且限制商业用途;样本部分源自既有基准错误案例,仍需保留来源、核对第三方图像与数据权利。公开样本可能包含人物、文本或联系方式等内容,处理和展示时应做隐私与敏感信息检查。

获取、校验与交付

咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。