Skip to main content

8 posts tagged with "ai"

View All Tags

Two flows: why a second, simpler model runs beside the primary

· 9 min read
Ivan Baha
Software Team Lead & Architect

The most expensive thing my agent can do is ask its smartest model a trivial question.

That sounds backwards – trivial questions are supposed to be cheap. A two-line summary, a yes/no judgement, a bit of text cleanup: milliseconds of honest work for a competent model. But Azek runs on one machine, and on one machine the flagship model's attention is the scarcest resource in the system. Every small job it handles personally gets paid for twice: once in the GPU it occupies, and once in a currency most people never see billed.

So Azek – the self-hosted agent from the first post – runs two models side by side, on purpose. The big one talks to me. The small one does everything else. This post is about the second lane: why it exists, what drives on it, and why the split turned out to be one of the best decisions in the whole project.

Why I'm building my own AI agent (and why it runs on my hardware)

· 9 min read
Ivan Baha
Software Team Lead & Architect

Last month my AI agent spent eight minutes turning one email into a calendar event. Five of those minutes were pure waste – the model fumbling with tool parameters, failing validation, retrying, fumbling again. I sat there watching tool_input_invalid errors scroll past, produced by software I own completely, running on hardware I paid for, doing a task any cloud assistant finishes before you've put the phone down.

I'd do it all again tomorrow. This post is about why – and it's the opening move of a series about what it actually takes. (Part 2 looks at why a second, simpler model runs beside the primary.)

Because the obvious question deserves a straight answer. It's 2026. Capable AI agents are a commodity: sign up, connect Gmail, press "allow," done. Why would anyone build one from scratch and insist it runs on a laptop?

First touch to Claude Fable 5 - A Capricious Vibe-Coding Frontier

· 7 min read
Ivan Baha
Software Team Lead & Architect

With the general availability of Anthropic’s new "Mythos-class" reasoning model, Claude Fable 5, the developer ecosystem has been flooded with benchmarks hailing its long-horizon autonomous capabilities. But synthetic leaderboards rarely paint an accurate picture of day-to-day repository engineering.
To see how Fable 5 holds up under real conditions, I ran an end-to-end security and logic audit on a real-world repository – nestjs-env-getter, a zero-runtime-dependency configuration manager for NestJS applications. I pitted Fable 5 (Max Effort) against the seasoned veteran Opus 4.8 (Max Effort), using Sonnet 4.6 (High Effort) as the execution muscle to see which model writes a better project blueprint for downstream automation.
The results exposed a massive paradigm shift in how frontier models approach codebase context, safety guardrails, and downstream delegation. Here is the comprehensive breakdown of my first brief testing.

Claude Fable 5: A Breakthrough in Cybersecurity or a Comedy of Guardrails?

· 4 min read
Ivan Baha
Software Team Lead & Architect

Two days ago, Anthropic launched Claude Fable 5, pitching it to the public as a breakthrough in cybersecurity capability. Naturally, I put it to the test on some of my own codebases: a lightweight lib for NestJS apps (nestjs-env-getter) and a highly complex, 40k+ LOC private AI agent orchestrator. The results weren't just disappointing – they exposed a structural comedy of guardrails that completely breaks the promise of AI-driven defence for regular engineers.

Enterprise AI Compliance: Steerings Over Skills

· 5 min read
Ivan Baha
Software Team Lead & Architect

Engineering teams frequently misconfigure AI coding agents (Copilot, Kiro, Claude, Codex, and others) by relying exclusively on on-demand skills — callable procedures and tool descriptions defined in an agent's tool registry (such as those conforming to the agentskills.io standard) — to enforce strict project rules. This approach fundamentally misunderstands the architecture of foundation models. When compliance mandates are embedded within dynamic tool descriptions, they are subject to conditional execution and instructional distraction. To prevent agents from silently violating security baselines and tooling mandates, organisations must implement a dual-layer architecture: static steerings for immutable boundaries and dynamic skills for isolated execution.

The AI-Native Team Workspace: Solving the Multi-Repo Context Crisis

· 4 min read
Ivan Baha
Software Team Lead & Architect

Engineering teams are hitting a wall with modern AI coding agents. Tools like GitHub Copilot Workspace, Cursor, and Claude Code are incredibly capable, but they encounter a severe structural limitation in enterprise environments: they are blind outside their immediate repository.

If your architecture consists of a React frontend in one repo, Node.js microservices in another, and Terraform manifests in a third, an AI agent operating in the frontend cannot trace a failing API call down to the database schema. It lacks the system-wide context required to make accurate, architectural-level contributions.

Emergent Creativity: An Architectural View on AI Consciousness and Deception

· 5 min read
Ivan Baha
Software Team Lead & Architect

Early artificial intelligence research, which started in the early 1950s, split into two distinct architectural paradigms. The first was a logic-inspired approach that attempted to hard-code intelligence using symbolic expressions and predefined rules. The second, biologically inspired approach posited that intelligence is fundamentally rooted in learning through networks of simulated brain cells. Rather than writing explicit logic, this architecture focused on enabling a system to learn by recognizing patterns and making analogies. It was inspired by research into how our brain works, realizing that biological networks are highly effective at finding analogies and patterns, and then using them to recreate or recognize information.