AI/ML, robotics & automation and computer vision project notes by Kukil Kashyap Borgohain

Inside Edge0-35B: Running a 35B Sparse MoE in Under 3 GiB RAM via SSD Offload and Prerouter Speculation

Inside Edge0-35B: Running a 35B Sparse MoE in Under 3 GiB RAM via SSD Offload and Prerouter Speculation
AI/ML

A deep dive into Edge0-35B-A3B's streaming MoE architecture, 2.9 GiB active memory footprint, double-shift prerouter speculation, and Recover-LoRA distillation on Apple Silicon.

15 min read
Read Full Article

Latest Blog Posts

Explore my latest thoughts and tutorials

STAY CONNECTED

Subscribe to My Newsletter

Project write-ups, model architectures, and embedded automation tutorials straight to your inbox.