Preview: PrismML's Bonsai 27B, open-weight MoEs, and local agent control planes.

Hey — Running larger models locally has always been a game of hardware compromise. This week, we saw some clever software and quantization tricks that suggest we might be closer to bypassing the VRAM bottleneck entirely.

🔥 This Week's Big One

Bonsai 27B by PrismML

Bonsai 27B is a 27-billion parameter multimodal model (derived from Qwen3.6-27B) compressed down to 1-bit (3.9 GB) and ternary (5.9–7.2 GB) variants that can run directly on consumer devices. By running locally in-browser via custom WebGPU kernels, it demonstrates that near-FP16 precision is achievable on a 10GB memory footprint, making phone-based local inference of large models a practical reality. https://huggingface.co/prism-ml/Bonsai-27B-gguf

🧠 Model Releases

🛠 LM Studio / Ollama

🤖 Agents & Frameworks

📦 Tools & Repos

⚡ Quick Hits


Until next week, keep your inference local and your cache warm.

— the LocalAILab team