Back to timeline

Fri, July 1003:31Model/APIAPI & pricingInfra & costAI hardwareAPI & pricing guide

NVIDIA Releases Nemotron Labs 3 Puzzle 75B A9B: Compressed MoE Model with 2.03x Throughput

Decision Brief

What changedNVIDIA unveiled the Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of the Nemotron-3-Super.
Why it mattersUsing hardware-aware structured compression and knowledge distillation, it reduces parameters from 120.7B to 75.3B while achieving 2.03x throughput on a single 8×B200 node, significantly lowering deployment costs and improving concurrency.
Who should careTeams building on model APIs, Inference / infra teams
Affected stackNVIDIA
Source confidenceMedium · Reliable media or first-hand reporting

NVIDIA released the Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of the Nemotron-3-Super model. This model uses an iterative Puzzle approach, alternating between hardware-aware structured compression and short-term knowledge distillation recovery phases. Total parameters are reduced from 120.7B to 75.3B, and active parameters from 12.8B to 9.3B. On a single 8×B200 node, the model achieves 2.03x the throughput of the original, delivering 100 tok/s per user. On a single H100, concurrency for 1M tokens increased from 1 to 8 requests.

Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.

Sources

  • MarkTechPost

    Fast research-paper and ML tooling summaries, useful for infra and agent updates.

  • MarkTechPost

Related intel

留言

登入后即可留言,和其他 builder 交换实测心得。

还没有留言,抢头香。