Back to timeline

Thu, July 916:47Model/APIAPI & pricingInfra & costAI hardwareAPI & pricing guide

NVIDIA Releases Compressed MoE Model Nemotron-Labs-3-Puzzle-75B-A9B, Boosting Throughput 2.03x

Decision Brief

What changedNVIDIA launches Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super.
Why it mattersReducing parameters from 120.7B to 75.3B with only 9.3B active, doubling single-node throughput, enabling teams to cut hardware costs significantly.
Who should careTeams building on model APIs, Inference / infra teams
Affected stackNVIDIA
Source confidenceMedium · Reliable media or first-hand reporting

NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE model from Nemotron-3-Super. Using iterative Puzzle method alternating hardware-aware structure compression and short knowledge distillation recovery, total parameters dropped from 120.7B (active 12.8B) to 75.3B (active 9.3B). On a single 8x B200 node, total throughput reached 2.03x the original Super, while maintaining 100 tok/s per user. On a single H100 GPU, concurrent 1M token requests increased from 1 to 8. For teams deploying large models on NVIDIA GPUs, this means supporting more users or higher concurrency on the same hardware, or halving the number of GPUs needed for equivalent performance.

Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.

Sources

  • MarkTechPost

    Fast research-paper and ML tooling summaries, useful for infra and agent updates.

  • MarkTechPost

Related intel

留言

登入后即可留言,和其他 builder 交换实测心得。

还没有留言,抢头香。