Preview: Liquid AI's LFM2.5 2.6B, Qwen 3.8 open weights, full-duplex voice, and EU AI Act tools.

Hey — As local hardware constraints force us to think beyond VRAM stacking, non-transformer architectures and deterministic safety gates are quietly redefining how we run agents on-device.

🔥 This Week's Big One

Liquid LFM2.5-2.6B: Non-Transformer Edge Intelligence

Liquid AI released LFM2.5-2.6B, a 2.69B parameter hybrid model that trades standard Transformer self-attention for linear recurrence and state-space layers. Featuring a 128K context window and requiring under 2.5 GB of RAM when quantized, it achieves speeds up to 220 tokens/sec on Apple Silicon and runs comfortably on mobile hardware. It proves that architectural efficiency—rather than raw VRAM capacity—is becoming the primary lever for high-throughput on-device inference. Hugging Face

🧠 Model Releases

🛠 LM Studio / Ollama

🤖 Agents & Frameworks

📦 Tools & Repos

⚡ Quick Hits


Keep your models local and your execution loops grounded.

— Himanshu