NVIDIA 发布压缩混合 MoE 模型 Nemotron-Labs-3-Puzzle-75B-A9B,吞吐量提升 2.03 倍
Decision Brief
What changedNVIDIA 推出 Nemotron-Labs-3-Puzzle-75B-A9B,是 Nemotron-3-Super 的压缩变体。
Why it matters参数量从 120.7B 降至 75.3B 而活跃参数仅 9.3B,单节点吞吐量翻倍,运行该模型的团队可大幅降低硬件成本。
Who should careTeams building on model APIs, Inference / infra teams
Affected stackNVIDIA
Source confidenceMedium · Reliable media or first-hand reporting
NVIDIA 发布 Nemotron-Labs-3-Puzzle-75B-A9B,这是 Nemotron-3-Super 的压缩混合 MoE 模型。通过迭代式 Puzzle 方法,交替进行硬件感知结构压缩和短知识蒸馏恢复阶段,模型总参数量从 120.7B(活跃 12.8B)降至 75.3B(活跃 9.3B)。 在单台 8x B200 节点上,该模型总吞吐量达到原版 Super 的 2.03 倍,同时保持每用户 100 tok/s 的输出速度。在单张 H100 GPU 上,1M token 并发请求数从 1 提升至 8。对于使用 NVIDIA GPU 部署大模型的团队,这意味着相同硬件能支持更多用户或更大并发,或为达到同等性能所需 GPU 数量减半。
Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.
Sources
- MarkTechPost
Fast research-paper and ML tooling summaries, useful for infra and agent updates.
- MarkTechPost
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。