模型

OpenAI gpt-oss-120b

openai/gpt-oss-120b

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 内容检查:AI 辅助整理并检查 许可证:Apache License 2.0 上游版本:b5c939de8f75

中文简介

gpt-oss-120b 是 OpenAI 发布方开放权重系列中的较大版本,固定 README 标注约 117B 总参数、5.1B 激活参数,面向通用和较高推理需求。它与 ChatGPT 在线产品是不同交付形态;下载权重不会获得 ChatGPT 的托管工具、账户功能或服务保障。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

版本定位:gpt-oss-120b 是 OpenAI 开放权重系列的大版本,适合已有高显存 GPU、需要更高容量推理的团队。它不是 ChatGPT 在线服务的离线副本。

选择建议:固定卡片列出约 117B 总参数、5.1B 激活参数,并给出 MXFP4 单 80 GB GPU 路线。仍应先用 20b 验证 Harmony 格式、任务指标和安全策略,再决定是否增加存储与硬件成本。

文件与边界:仓库同时保存多种权重布局,整仓库约 182 GB,不代表运行时必须重复加载全部格式。Apache 2.0 不替代输出安全和数据合规审查;本站只提供版本核对与合规交付咨询。

已有简体中文译文 · Codex 基于固定版本上游 README 编写;待人工复核 · 2026-08-07 23:24

Try gpt-oss · Guides · Model card · OpenAI blog Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: `gpt-oss-120b` — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) `gpt-oss-20b` — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. [!NOTE] This model card is dedicated to the larger `gpt-oss-120b` model. Check out `gpt-oss-20b` for the smaller model. Highlights **Permissive Apache 2.0 license:** Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment. **Configurable reasoning effort:** Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs. **Full chain-of-thought:** Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in outputs. It’s not intended to be shown to end users. **Fine-tunable:** Fully customize models to your specific use case through parameter fine-tuning. **Agentic capabilities:** Use the models’ native capabilities for function calling, web browsing, Python code execution, and Structured Outputs. **MXFP4 quantization:** The models were post-trained with MXFP4 quantization of the MoE weights, making `gpt-oss-120b` run on a single 80GB GPU (like NVIDIA H100 or AMD MI300X) and the `gpt-oss-20b` model run within 16GB of memory. All evals were performed with the same MXFP4 quantization. Inference examples Transformers You can use `gpt-oss-120b` and `gpt-oss-20b` with Transformers. If you use the Transformers chat template, it will automatical

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

适用场景

适合具备 80 GB 级 GPU 的团队评估较高容量本地推理、代理和专用知识任务,也适合与 20b 在质量、时延和成本上做同框架比较。应先确定是否真的需要 120b;许多 PoC 用 20b 更容易完成迭代与安全测试。

模型参数

固定版本:b5c939de8f754692c1647ca79fbf85e8c1e70f8a。Hugging Face 元数据参数量:120,412,337,472;固定 README:约 117B 总参数、5.1B 激活参数、MXFP4,并按 Harmony response format 使用。

文件说明

固定 SHA 下共有 37 个文件,包含 Transformers、original 和 Metal 相关权重布局,文件元数据合计约 182.32 GB,API used_storage 约 182.40 GB。单一部署路径无需机械获取所有重复格式,但选定格式的分片、索引、config、tokenizer 和 LICENSE 必须完整。

上游文件元数据

  • .gitattributes1.53 KB
  • chat_template.jinja16.35 KB
  • config.json2.04 KB
  • generation_config.json177 B
  • LICENSE11.09 KB
  • metal/model.bin60.76 GB
  • model-00000-of-00014.safetensors4.31 GB
  • model-00001-of-00014.safetensors3.83 GB
  • model-00002-of-00014.safetensors4.31 GB
  • model-00003-of-00014.safetensors3.83 GB
  • model-00004-of-00014.safetensors4.31 GB
  • model-00005-of-00014.safetensors3.83 GB

硬件建议

发布方固定 README 给出单张 80 GB GPU 的 MXFP4 运行目标。生产部署还需为 KV cache、并发、服务框架和故障冗余预留资源;普通 80 GB 卡能加载不等于满足目标吞吐。微调、长上下文和多用户服务应另做容量测试。

注意事项

模型使用 Harmony 格式,需选择明确支持 gpt-oss 的推理实现。发布方所称单 80 GB GPU 运行基于 MXFP4 路径,并非任意精度、任意上下文和任意并发都适用。工具调用、完整推理内容、输出合规和第三方数据仍需业务侧治理。

获取、校验与交付

咨询此资源时只需发送本页链接或资源准确全称。橙子AI科技会继续核对版本、文件与类型、README资料卡、许可证和访问条件,并在合法访问权限、许可证及平台规则允许的前提下,协助海内外下载、完整性校验及网盘或硬盘交付;本官网本身不托管或下载资源文件。

第三方资源声明

本页面为橙子AI科技基于固定版本上游卡片整理的中文信息与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。