Unsloth v0.1.900-beta adds local Laya decision-model serving through a Decision API | Tool Update
Unsloth published v0.1.900 beta on 28 Sep 2026, adding support for running and serving Decision Models like Laya locally through a Decision API enabled under S…
Published on MyPrivateClaw
Sep 29, 2026, 1:19 AM UTC
Coverage date
Sep 28, 2026
Last updated
Sep 29, 2026, 1:19 AM UTC
News summary
Unsloth published v0.1.900 beta on 28 Sep 2026 (release timestamped 28 Sep 14:33), adding support for running and serving Decision Models like Laya locally, exposed through a Decision API enabled under Settings API with CPU or GPU controls. Laya integrates with the TypeSafe SDK through a Jev compatible /v1/systemone endpoint, and runs natively on MLX for Apple Silicon when GPU is selected. For Apple Silicon + MLX, v0.1.900 beta added TurboQuant KV cache and KV cache quantization for sliding window models such as Gemma 4, plus batched serving and grammar constrained structured outputs through response format. The release also carries vendor claimed speedups: Unsloth claims LTX 2.3 clips are 4.5x faster with distilled sampling, compile fixes, and hosted FP8 weights, and that image and video VAE optimizations deliver 1.7 6.3x faster decoding. It further claims MiniMax H3's first render is…