MyPrivateClaw

NVIDIA Reports Nemotron 3 Ultra Outperforms Kimi K2.6 and GLM 5.2 on Agentic RTL Coding Benchmark

A July 26, 2026 performance report from NVIDIA shows Nemotron 3 Ultra achieved 97.1% average pass rate on CVDP benchmark using ACE RTL agent framework, outperf…

Published on MyPrivateClaw

Jul 27, 2026, 4:17 PM UTC

Coverage date

Jul 27, 2026

Last updated

Jul 27, 2026, 4:17 PM UTC

News summary

Nemotron 3 Ultra features 550B total parameters with 55B active parameters using LatentMoE hybrid Mamba 2 + MoE + Attention architecture with Multi Token Prediction, available as BF16 weights and NVFP4 quantization on Hugging Face. The model supports up to 1M token context length and requires a minimum of 8x GB200/B200/GB300/B300 or 16x H100 GPUs for deployment. On the CVDP benchmark using ACE RTL with Nemotron 3 Ultra as the model backbone, a 97.1% average pass rate was achieved across nine agentic RTL task categories including RTL code completion, spec to code, modification, module reuse, linting/QoR improvement, testbench stimulus generation, testbench checker generation, assertion generation, and debugging/bug fixing. Nemotron 3 Ultra used an average of 6,629 tokens per iteration on CVDP tasks with ACE RTL, approximately 28% fewer tokens than GLM 5.2 (9,156) and 71% fewer than Kimi…