NVIDIA 发布 Nemotron Labs 3 Puzzle 75B A9B:压缩混合 MoE 模型,吞吐量提升 2.03 倍
Decision Brief
What changedNVIDIA 发布了 Nemotron-Labs-3-Puzzle-75B-A9B,这是 Nemotron-3-Super 的压缩变体。
Why it matters该模型通过硬件感知结构压缩和知识蒸馏,将参数从 120.7B 降至 75.3B,同时在单节点 8×B200 上实现 2.03 倍吞吐量,大幅降低部署成本并提升并发能力。
Who should careTeams building on model APIs, Inference / infra teams
Affected stackNVIDIA
Source confidenceMedium · Reliable media or first-hand reporting
NVIDIA 发布了 Nemotron-Labs-3-Puzzle-75B-A9B,这是 Nemotron-3-Super 的压缩变体。该模型采用迭代 Puzzle 方法,交替进行硬件感知结构压缩和短期知识蒸馏恢复阶段。总参数从 120.7B 降至 75.3B,激活参数从 12.8B 降至 9.3B。在单个 8×B200 节点上,该模型实现 2.03 倍于原始模型的吞吐量,每个用户达到 100 tok/s。在单块 H100 上,1M token 并发从 1 个请求提升至 8 个。
Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.
Sources
- MarkTechPost
Fast research-paper and ML tooling summaries, useful for infra and agent updates.
- MarkTechPost
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。