
AI features create tech debt faster than most features, and it's rarely the model's fault. The debt comes from how the model gets wired in: prompts scattered through the codebase, models called from everywhere, versions nobody pinned, and an agent asked to do everything.
None of this shows up in the first month. It shows up when you try to change something and can't tell what will break. Here is where the debt comes from, and the decisions that prevent it.
When every part of the product calls a model directly, every change to the model is a change to everything. Put AI behind its own boundary instead.
On Nooma, the AI practice companion we designed and built with O-HR, the platform API carries auth, tenancy, billing and the catalogue, and a separate AI service runs the agent runtime. The rest of the product asks the AI service for work and doesn't care how it gets done. That separation is what lets you swap a model, change a prompt or add an agent without touching billing.
Prompts written inline in application code are the most common source of AI debt. Nobody can see them all at once, nobody knows which version is live, and nobody can test a change properly.
Keep prompts in one place, version them like any other configuration, and tie each version to the evaluation set that proves it works. We cover evaluation in stop shipping AI demos.
"Latest" is not a version. A provider's model update can change outputs overnight, and embeddings from a different model aren't compatible with the ones you already stored. On Nooma, embeddings run on a version-pinned model for exactly this reason. Upgrading becomes a deliberate change you test, not a surprise you debug.
One agent that does everything is hard to test, hard to reason about and hard to govern. Nooma instead runs purpose-built agents, each scoped to a single task with its own knowledge base, forms and decision gates. Smaller scope means smaller failures, clearer evaluation and a much easier conversation when someone asks what the AI is allowed to do.
Open-ended retrieval is debt in disguise. The more the AI can reach, the harder it is to explain an answer or guarantee it didn't come from somewhere it shouldn't. Define the sources the feature works from and keep it inside them. On Nooma that's a governed corpus of Australian employment law and regulation, lawyer-verified templates and the organisation's own documents, with no open web.
Model costs scale with usage, which means your most successful feature can become your most expensive one. Decide early:
Every model provider, vector store and logging tool is a place your data goes. List them, check their terms on retention and training, and review the list when anything changes. On Nooma every subprocessor is disclosed and reviewed annually. It's dull work that pays for itself the first time a customer's security team asks. If personal information is involved, check your Privacy Act obligations before you build, not after.
Rarely, and only when the craft is the product. Most features are well served by a hosted model under commercial terms, wrapped in good architecture. OLi Tutor was the exception: the tutoring behaviour is the reason people choose it, so we built and trained the system that carries it. The test is simple. If a competitor could match your AI with a prompt, you don't need your own model. If they couldn't, you might.
If you're adding AI to a product that already exists, Strategy & Audit is a fixed-scope review that covers where AI fits and what the codebase needs first. Or see how we approach AI properly embedded.
Book a free 30-minute call. We'll talk through what you're working on, what we'd do, and whether we should partner. No pitch deck, no PDF brochure.