MyPrivateClaw

ISTA-DASLab publishes non-uniform GSQ-RCO quantizations of Qwen3.8-Flash-Next in 66.4 GB, 68.0 GB, and 75.8 GB builds running unmodified in llama.cpp, Ollama, and LM Studio

The ISTA DASLab Hugging Face workspace published non uniform GSQ RCO quantizations of Qwen3.8 Flash Next, a 512 expert MoE model, in three sizes at 66.4 GB, 68…

Published on MyPrivateClaw

Sep 16, 2026, 11:17 PM UTC

Coverage date

Sep 16, 2026

Last updated

Sep 16, 2026, 11:17 PM UTC

News summary

Published by the ISTA DASLab Hugging Face workspace, the non uniform GSQ RCO quantizations of Qwen3.8 Flash Next, a 512 expert MoE model, arrive in three sizes at 66.4 GB, 68.0 GB, and 75.8 GB and run unmodified in llama.cpp, Ollama, and LM Studio. The Q2 0 build delivers 3.4x the prompt throughput and 1.9x lower end to end latency than IQ2 XS, which the lab attributes to quantization format rather than file size, since the two files differ by only 1.6 GB. These llama.cpp measurements over 55 prompts are reported by the lab. At 3.00 bpw, the IQ3 XXS build matches the BF16 base exactly on AIME25 (100.00) at 75.8 GB, about one fifth of the 354 GB BF16 footprint, and trails by 0.51 on GPQA Diamond and 1.14 on LiveCodeBench v6. Those numbers are relative to the BF16 base model and are reported numbers. The quantized weights inherit the license of the base model Qwen3.8 Flash Next, while the…