Two flows: why a second, simpler model runs beside the primary
The most expensive thing my agent can do is ask its smartest model a trivial question.
That sounds backwards – trivial questions are supposed to be cheap. A two-line summary, a yes/no judgement, a bit of text cleanup: milliseconds of honest work for a competent model. But Azek runs on one machine, and on one machine the flagship model's attention is the scarcest resource in the system. Every small job it handles personally gets paid for twice: once in the GPU it occupies, and once in a currency most people never see billed.
So Azek – the self-hosted agent from the first post – runs two models side by side, on purpose. The big one talks to me. The small one does everything else. This post is about the second lane: why it exists, what drives on it, and why the split turned out to be one of the best decisions in the whole project.

