Every major consulting firm offers its own prebuilt architectural framework, model, or reference architecture for generative AI and agentic AI. The pitch is nearly identical from firm to firm: “We’ve done this before, we’ve packaged the patterns, and if you start from our blueprint, you’ll get down the path faster.” It’s an age-old consulting technique. Get the client interested in an out-of-the-box solution, then bill the hours to customize it. The client, as is so often the case, ends up holding the short end of the stick. To be clear about what’s out there right now, a short list of the players is warranted:
- McKinsey publishes an open reference architecture for scaling generative AI, built around reusable application patterns such as knowledge management, customer chatbots, and agentic workflows, plus shared services for retrieval-augmented generation (RAG), embedding, reranking, and intent classification, tied together with infrastructure as code and policy as code.
- Deloitte has published an AI agent architecture, a multi-agent systems framework, and an “agentic enterprise” blueprint that positions autonomous agents as the enterprise’s operating logic.
- Capgemini launched its Resonance AI Framework in July 2025, backed by the RAISE gallery of generative AI and agent patterns, promising to translate AI ambition into measurable business impact.
- Infosys offers a five-layer reference architecture for mature enterprise AI as part of its Topaz push, with each layer having its own personas, interfaces, and deployment model.
- TCS sells AI WisdomNext, an aggregation and orchestration platform that enables model swapping and deploys prebuilt agents without vendor lock-in.
- Accenture promotes an AI architecture for the intelligent enterprise, an “AI brain” centered on agent orchestration and decision intelligence.
One important caveat: Each framework serves a different purpose and audience. Others exist, and more arrive every quarter. I list these only as examples, not for detailed review; read them and draw your own conclusions. My concern isn’t any single framework, but the pattern of treating one as your starting point.
I understand exactly why these frameworks were created. They’re tools for selling consulting hours. A prebuilt framework gives a sales team something concrete to present to a CIO. It tells an attractive story, and, in fairness, some of the thinking in these documents is genuinely good. The problem isn’t that a prebuilt framework is wrong. The problem is what happens when you use them.
A good answer or the best one?
When you build any AI system, you’re creating a system that leverages dynamic behavior as a functional advantage for the application. That behavior depends entirely on your requirements. For any given AI problem, roughly 5,000 permutations of architecture and technology solutions will work. But only one is optimized for your requirements. The architect and architecture team must figure out which one it is, and the only way to do that is to start with the requirements and work toward the architecture, not the other way around.
A prebuilt framework muddies the waters. Both the consultants and the clients end up anchored to a static idea or concept, and then they try to force their problem domain into it. That’s backward. In every instance I’ve seen in the past two years alone, this dynamic ends with companies deploying technology patterns that simply aren’t needed. Vector databases where a search index would do. Multi-agent orchestration where a single model and good prompts will suffice. Layer upon layer of governance and middleware that no one asked for, justified because “it’s part of the framework.”
The cost of someone else’s plan
The result is a hugely underoptimized architecture that, in my direct experience, costs somewhere between 10 and 20 times more to operate than a fully optimized architecture would. And I mean fully optimized in the sense of an architecture where the team started from zero, incorporated their own requirements, and worked their way to the finish. No framework to distract them. No model. No architectural conceptualization gimmicks. They just solved the problem.
Think about what that multiplier means in real money. A monthly inference bill that should be $50,000 becomes $500,000 to $1,000,000. Not because the technology is expensive, but because the architecture was chosen before the requirements were understood, and every downstream decision inherited that mistake.
What you should do instead
Focus on requirements, procedures, and desired outcomes to shape your architecture. Spend the time up front to understand what dynamic behavior the application actually needs, what data it touches, what latency and cost envelopes it must live inside, and what failure looks like.
Work through ideas and concepts honestly, with no vendor intellectual property on the table. When you do this, you’ll
[…]
Content was trimmed to protect the source. Please visit the original article for the full text.
Read the original article:
