Tsinghua's Puro-2B Base Model Released Under Apache License 2.0
The thu pacman group at Tsinghua University has released Puro 2B, a 2B parameter dense causal language model pretrained from scratch on consumer grade NVIDIA R…
Published on MyPrivateClaw
Aug 28, 2026, 11:17 PM UTC
Coverage date
Aug 28, 2026
Last updated
Aug 28, 2026, 11:17 PM UTC
News summary
Puro 2B uses a Qwen3 1.7B compatible architecture with untied input and output embeddings, 28 transformer layers, 2,048 hidden size, blockwise FP8 (E4M3) for Transformer linear layer GEMMs, MuonH optimizer with hyperball constraints, and a context length of 4,096 tokens. The canonical Puro 2B Base model was trained at a measured accelerator cost of $6,891 (using 22,514 active GPU hours; accelerator only reproduction estimate; excludes data acquisition, preprocessing, evaluation, and research labor), and the report derives a Puro Cost Scaling Law estimating approximately $4.4K is sufficient to reach performance comparable to Qwen2 1.5B (recipe and hardware specific empirical estimate without uncertainty intervals). On a 15 task base model evaluation protocol (same deterministic OpenCompass evaluation pipeline; greedy decoding for generation tasks, fixed token likelihood ranking for multi…