模型

nemotron-3.5-asr-streaming-0.6b

nvidia/nemotron-3.5-asr-streaming-0.6b

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:other 上游版本:f3d333391852

中文简介

Nemotron 3.5 ASR Streaming 0.6B 是 NVIDIA 发布的流式多语言自动语音识别模型。标签显示其采用缓存感知 FastConformer-RNNT 架构,面向低延迟语音转写;支持语言、采样率、缓存策略、端点检测与许可证条件应以上游模型卡为准。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Nemotron 3.5 ASR Streaming 0.6B 是 NVIDIA 发布的流式多语言自动语音识别模型。标签显示其采用缓存感知 FastConformer-RNNT 架构,面向低延迟语音转写;支持语言、采样率、缓存策略、端点检测与许可证条件应以上游模型卡为准。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Nemotron 3.5 ASR h1, h2, h3, h4, h5, h6 { color: #76b900; /* NVIDIA green */ font-weight: 700; } hr { border: none; border-top: 1px solid #e5e7eb; margin: 2rem 0; } /* Improve list spacing */ ul, ol { margin-top: 0.5rem; margin-bottom: 0.5rem; } /* Badge alignment consistency */ img { display: inline; vertical-align: middle; }     [!Note] This model is the multilingual extension of nvidia/nemotron-speech-streaming-en-0.6b, adding language-ID prompt conditioning to support transcription across **40 language-locales** from a single model. Nemotron 3.5 ASR** is a multilingual, streaming Automatic Speech Recognition (ASR) model engineered to deliver high-quality multilingual transcription across both low-latency streaming and high-throughput batch workloads. Developed by NVIDIA, this 600M parameter model transcribes speech into text with native support for punctuation and capitalization, and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 320ms, 560ms, and 1120ms. By leveraging a state-of-the-art **Cache-Aware FastConformer-RNNT** architecture, the model eliminates redundant overlapping computations common in traditional "buffered" streaming. This allows it to process only new audio chunks while reusing cached encoder context, significantly improving computational efficiency and minimizing end-to-end delay without sacrificing accuracy. It was trained on a massive ASR dataset and is engineered to perform across diverse and challenging acoustic conditions. This model is ready for commercial use. Release Date Hugging Face [06/04/2026] via https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b Why Choose Nemotron 3.5 ASR? 🌍 **Single Multilingual Model:** Transcribes 40 language-locales from one model through language-ID prompt conditioning, with optional automatic language detection. ⚡ **Native Streaming Architecture:** Cache-aware design enables efficient processing of continuous audio streams, designed and optimized for lo

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes1.57 KB
  • arch_slide10.png85.78 KB
  • avg_wer_summary.png63.98 KB
  • bias.md2.05 KB
  • config.json1.34 KB
  • explainability.md2.33 KB
  • fleurs_langid_vs_auto.png81.88 KB
  • fleurs_wer_vs_chunk_size.png90.07 KB
  • generation_config.json193 B
  • latency_vs_parallel.png135.91 KB
  • model.safetensors2.38 GB
  • model_architecture.png147.06 KB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。