NVIDIA Releases Compressed MoE Model Nemotron-Labs-3-Puzzle-75B-A9B, Boosting Throughput 2.03x
Decision Brief
NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE model from Nemotron-3-Super. Using iterative Puzzle method alternating hardware-aware structure compression and short knowledge distillation recovery, total parameters dropped from 120.7B (active 12.8B) to 75.3B (active 9.3B). On a single 8x B200 node, total throughput reached 2.03x the original Super, while maintaining 100 tok/s per user. On a single H100 GPU, concurrent 1M token requests increased from 1 to 8. For teams deploying large models on NVIDIA GPUs, this means supporting more users or higher concurrency on the same hardware, or halving the number of GPUs needed for equivalent performance.
Sources
- MarkTechPost
Fast research-paper and ML tooling summaries, useful for infra and agent updates.
- MarkTechPost
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。