Notes on building AI that works.
Practical writing on what holds up in production, and what doesn't.
Opinions are not my employer's.
What Intelligence Should Cost
Most people buy intelligence the same way they buy a car. Bigger, and hope the next one is cheaper. We run a small lab that asks what intelligence actually costs, and whether you can buy the same capability for less.
Read →Switching From Ollama to llama.cpp For Local Inference
Local agentic coding is finally good enough to start taking architecture and performance more seriously. By switching from Ollama to llama.cpp, everything has improved for me. Speed, tool-calling, reliability...everything.
Read →The "unlimited" AI plan was always going to end
The big AI companies are shifting from flat "unlimited" subscriptions toward paying for every token you use. Here is why that was inevitable, and why running your own models and tools is the best protection against being priced out.
Read →