
Inside Qwen3.8-27B: Hybrid DeltaNet Architecture, 256k Context, and 24GB Local Serving
A deep dive into Qwen3.8-27B's hybrid Gated DeltaNet attention, 75% KV cache reduction, 256k native context, and deployment trade-offs on 24GB GPUs.

A deep dive into Edge0-35B-A3B's streaming MoE architecture, 2.9 GiB active memory footprint, double-shift prerouter speculation, and Recover-LoRA distillation on Apple Silicon.

Explore my latest thoughts and tutorials

A deep dive into Qwen3.8-27B's hybrid Gated DeltaNet attention, 75% KV cache reduction, 256k native context, and deployment trade-offs on 24GB GPUs.

CadCore translates natural language prompts into verified parametric 3D CAD models (.step, .stl), 2D technical drawings (.svg), and interactive Three.js viewers using closed-loop execution and error repair.

A beginner's guide to the open-weight voice LLM ecosystem on Hugging Face in 2026. Covers neural audio codecs, full-duplex speech models, VRAM sizing, and local serving.

How to safely roll out candidate models to live voice telephony traffic. Part 2 covers bot-to-bot user simulation, dark traffic shadowing, canary routing, and multi-model fallbacks.

How to survive foundation model deprecations in real-time voice bots. Part 1 covers failure modes, TTFS latency budgets, phonetic drift, and declarative CI/CD testing.

Explore the architecture and real-world performance of NVIDIA's new LocateAnything-3B, featuring its groundbreaking Parallel Box Decoding for high-speed object localization.
Project write-ups, technical deep dives, and working notes. Everything here was built, tested, or broken by hand.
LLM fine-tuning, local inference, model reviews, GenAI application architecture, and running open-weight models.
Object detection, OpenCV projects, YOLO deployments, vision model training, and edge AI on constrained hardware.
ESP32 and Arduino projects, embedded AI, hardware integration, sensor wiring, and IoT application builds.
The math and theory behind machine learning — probability, Bayesian thinking, kernel methods, and model internals explained clearly.
Workflow automation, scripting, process optimization, and developer tooling for cutting down repetitive work.
Ender 3 setup guides, slicer settings, bed calibration, print failure detection, and monitoring dashboards.
Project write-ups, model architectures, and embedded automation tutorials straight to your inbox.