Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
From Retrieval to Governance — The Architecture Shift Replacing Naive RAG
Latest   Machine Learning

From Retrieval to Governance — The Architecture Shift Replacing Naive RAG

Last Updated on October 6, 2026 by Editorial Team

Author(s): Sandeep Chaudhary

Originally published on Towards AI.

From Retrieval to Governance — The Architecture Shift Replacing Naive RAG

Stop asking “How do we retrieve more chunks?” Start asking “How do we govern and structure what the model sees?”

From Retrieval to Governance — The Architecture Shift Replacing Naive RAG

Naive RAG is not failing because the idea was wrong.

It is failing because the implementation stopped at the easy part — embedding chunks and running cosine similarity — and called it an architecture. The hard part is governance, structure, freshness, and determinism. The hard part is what production enterprise workloads actually require.

The industry is moving. The teams that are getting production AI right are not optimising their chunk sizes. They are replacing the retrieval mindset entirely — building governed, deterministic context layers that control precisely what the model sees at inference time, at what cost, with what latency, and with what audit trail.

That shift has a name. Context Engineering. And it changes the architecture significantly.

The Illusion of Vector Search

Here is what naive chunking actually does to your enterprise knowledge.

You take a 200-page legal contract. You split it into 512-token chunks. You embed each chunk and store it in a vector database. A user asks a question. You run cosine similarity against the query embedding. You retrieve the top-K chunks. You stuff them into a prompt.

What the model sees is a collection of decontextualised text fragments with no structural relationship to each other, no information about what came before or after each fragment, no cross-document dependencies, and no indication of which clause supersedes which.

The contract had structure. The chunking destroyed it. The model now has to reconstruct the semantic relationships that the chunking broke — from fragments that were never designed to be read in isolation.

This is semantic fragmentation. And it is not a tuning problem. It is an architectural problem with the naive approach.

The same failure mode appears across every enterprise domain where structure matters. A codebase chunked at 512 tokens loses the import graph, the class hierarchy, the function signature context. A technical specification loses the dependency chain between requirements. A policy document loses the precedence relationships between clauses.

Vector similarity finds tokens that look like the query. It does not find the context that is needed to reason correctly about the query. For simple lookup tasks, that gap is manageable. For complex enterprise reasoning — regulatory interpretation, code generation across a codebase, contract analysis, technical architecture questions — it is where accuracy stalls.

Anatomy of Failure: Why Naive RAG Stalls in Production

I have seen this pattern enough times to describe it precisely.

A team builds a RAG system. It works well on the demo dataset. It works reasonably well on the test set. It goes into production and accuracy on complex enterprise queries caps out somewhere between 50–65%. The team spends months trying to improve it — better chunking strategies, re-ranking, hybrid search, query reformulation. Each intervention adds marginal improvement. The ceiling does not move.

The ceiling is structural. It is not a retrieval problem. It is four problems that retrieval optimisation cannot fix.

Fatal Flaw 1: Lost Signal-to-Noise Ratio

Top-K retrieval optimises for semantic similarity to the query. It does not optimise for the completeness of the context needed to answer the query.

The result is retrieval sets that are semantically adjacent to the query but contextually incomplete. The relevant clause is retrieved but the definition section that gives it meaning is not. The function is retrieved but the interface it implements is not. The policy is retrieved but the exception that applies to the current case is not.

The model receives fragments that gesture toward the answer without providing the structural context required to reason to it correctly. Increasing K increases noise as much as it increases signal. The trade-off does not resolve — it gets worse as the knowledge base grows.

Fatal Flaw 2: Auditability and Lineage Gaps

In a regulated enterprise, every AI-assisted decision needs an audit trail. What information did the model see? What version of that information? Who had access to it? When was it last validated?

Vector stores cannot answer these questions natively. The chunk that was retrieved at query time is a fragment of a document at some embedding time. The original document may have since been updated, superseded, or revoked. The retrieval pipeline has no mechanism to express this. The audit log shows a cosine similarity score — not a document identifier, version, access policy, or freshness timestamp.

When the compliance team asks “what did the model see when it made this recommendation last Tuesday?” — a vector store with point-in-time retrieval semantics cannot reliably answer. This is not a tooling gap. It is a fundamental property of systems that trade governance for retrieval speed.

Fatal Flaw 3: No Freshness SLA

Chunks have no ownership. No data-quality gate. No version control. No expiry.

A chunk embedded six months ago reflects the state of the document six months ago. If the policy changed three months ago, the chunk still says the old thing. The retrieval system will confidently return it and the model will reason from stale information — with no indication to the user that anything is wrong.

Enterprise knowledge has a freshness requirement that varies by document type. Financial data may be stale within hours. Legal interpretations within weeks. Product specifications within days. A retrieval system with no freshness SLA cannot express these distinctions. It retrieves whatever was embedded most recently relative to the query — regardless of whether that information is currently valid.

Fatal Flaw 4: Lost in the Middle

This is the failure mode that is easiest to reproduce and hardest to fix within the naive RAG paradigm.

Long-context models perform significantly worse on information positioned in the middle of the context window than on information at the beginning or end. Passive chunk dumps — retrieving K chunks and concatenating them into the prompt — place the most relevant context at an unpredictable position in the prompt, often in the middle of a large retrieval block.

The model’s instruction-following capability degrades. The retrieved context is present but not effectively attended to. Accuracy drops on queries that require synthesising information from multiple retrieved chunks.

The mitigation — re-ranking to place the most relevant chunks at the beginning or end — helps at the margins. It does not fix the underlying problem that passive chunk dumps give the model no structural signal about how to weight and use what it is reading.

Paradigm Shift 1: Context Engineering

Context Engineering is the broader discipline of managing everything the model sees at inference time — not just retrieved text.

This includes the system prompt, the tool catalogue, the agent’s memory, the user’s role-based permissions, the governance constraints, the schema of the expected output, and the structured knowledge the model needs to reason correctly. All of it is context. All of it needs to be governed. The retrieval problem is a subset.

The shift in perspective is from: “how do we find the right text chunks?” to “how do we construct the optimal, governed context for this model at this inference moment?”

Write on Medium

Context Engineering has four layers.

Layer 1: Structural Curation

Replace raw text chunks with structured representations that preserve the relationships that chunking destroys.

  • Knowledge Graphs. Entities, relationships, and properties extracted from source documents and stored as a graph. A query against a knowledge graph retrieves not just the relevant entity but its relationships — which policies it references, which clauses it supersedes, which functions it calls. The structural information that chunking discards is the structural information that reasoning requires.
  • Abstract Syntax Trees (ASTs) for code Instead of chunking source files at token boundaries, parse them into their syntactic structure. Retrieve function definitions with their signatures, their dependencies, and their callers — not a fragment of their implementation. Code retrieval at the AST level produces context that the model can reason about as code rather than as text that happens to look like code.
  • Hierarchical Document Trees For long documents — legal contracts, technical specifications, regulatory filings — preserve the document hierarchy. Section → subsection → clause → sub-clause. Retrieval at the hierarchical level returns the clause with its ancestral context intact: which section it belongs to, which scope it operates within, which overriding provisions exist at higher levels

The shift from flat chunks to structured representations requires more investment at index time. It produces dramatically better context at inference time — context the model can actually reason across rather than fragments it has to reassemble.

Layer 2: Data Quality and Governance Gates

Before any context reaches the model, it passes through a governance layer.

  • Lineage tracking Every piece of context carries its provenance — source document, version, access tier, last validation date, owning team. The model’s context is auditable at the element level, not just the query level.
  • Freshness enforcement Context elements carry expiry policies aligned with their content type. Financial data with a 4-hour TTL. Policy documents with a 30-day TTL. Technical specifications validated on every release cycle. Stale context is not retrieved — it is flagged for refresh and excluded from inference until refreshed
  • Role-based context masking The context available to the model is filtered by the requesting user’s permissions before assembly. An analyst querying a compliance system sees public regulatory context. A compliance officer sees that plus restricted internal interpretation guidance. The model never sees context the user is not authorised to access — enforced at the context assembly layer, not by hoping the model respects a permission hint in the system prompt.
  • Column-level security for structured data When structured data is part of the context, field-level masking applies. The model sees the fields it is authorised to reason about and no others.

Layer 3: Dynamic Compression and Pruning

A million-token context window does not mean you should fill it. Token budget management is an active discipline, not a passive limit.

  • Context budget allocation The total token budget for an inference call is partitioned across its components: system prompt, tool schemas, retrieved knowledge, conversation history, and expected output headroom. Each component has a budget. When any component exceeds its budget, compression is applied — not by truncating arbitrarily, but by applying domain-aware summarisation or selective pruning.
  • Selective context pruning Not all retrieved context is equally relevant to the current query. A context scoring step evaluates each retrieved element against the query and the current agent state. Elements below the relevance threshold are excluded — not because they fail cosine similarity but because they are not needed for this specific reasoning task at this moment.
  • Prefix tuning for static context Parts of the context that do not change across a session — the system prompt, the tool catalogue, the organisation’s governance constraints — can be expressed as a frozen prefix that is processed once and cached. The dynamic suffix — the query, the conversation turn, the specific retrieved knowledge for this request — is the variable component. Separating static from dynamic context enables caching.

Layer 4: Protocol-Driven Context — MCP

The Model Context Protocol is a standardised interface for streaming structured, stateless context from tool servers to model clients.

Instead of building custom function-calling integrations for every data source — each with its own schema, authentication, error handling, and versioning — MCP provides a common protocol. Tool servers expose resources and tools via MCP. Model clients discover and invoke them through the same interface regardless of the underlying system.

  • The context engineering implication MCP servers can be designed as context servers — delivering pre-structured, governed context blocks directly to the model at inference time. Instead of the agent constructing its context from raw retrieval results, it requests a structured context package from a server that knows the domain, applies the governance rules, and returns a context block that is ready for model consumption.

The stateless nature of MCP means context is assembled fresh for each inference — no session state to manage, no context drift across turns, clean audit trail per request.

Paradigm Shift 2: Context Caching

Long-context windows changed the theoretical boundary of what a model can reason across. Context caching changed the economics of actually using those windows.

Without caching, a million-token context costs $15+ per call at current frontier model pricing, with a Time-to-First-Token (TTFT) measured in tens of seconds. At that cost and latency, million-token contexts are a capability claim, not a practical architecture choice for production enterprise systems.

Context caching changes the equation.

Prefix Caching Mechanics

Modern inference infrastructure supports KV (Key-Value) cache reuse. When a model processes a prompt, it computes attention Key-Value pairs for each token in the context. These KV pairs are the model’s internal representation of what it has read. They are computationally expensive to produce and — for static context that does not change between requests — identical every time they are computed.

Prefix caching stores the KV pairs for the static portion of the prompt. On subsequent requests that share the same static prefix, the KV pairs are loaded from cache rather than recomputed. The model skips straight to processing the dynamic suffix — the part that actually varies between requests.

  • Cost impact: Cached input tokens are typically priced at 10–20% of uncached token cost on frontier model APIs that support it. A system prompt and knowledge base that consume 100,000 tokens — computed once and cached — cost almost nothing on subsequent requests that reuse the cache.
  • Latency impact: TTFT on requests that hit a warm cache drops by 80–90%. The model is not reading 100,000 tokens of static context — it is reading the dynamic suffix only.

Cache-Aware System Design

Caching is not automatic. It requires deliberate architectural separation of static and dynamic context.

Cache Pre-Prefix (immutable, maximum cache lifetime):

  • System prompt and role definition
  • Tool schemas and capability catalogue
  • Organisation-level governance constraints
  • Domain knowledge bases that change infrequently

Dynamic Suffix (not cached, variable per request):

  • Current query or user message
  • Conversation history for this session
  • Request-specific retrieved context
  • Real-time data that changes between calls

Migration Playbook: From Naive RAG to Context Engineering

Step 1: Audit Your Chunks

Before changing anything, understand where the current system is losing information.

Step 2: Standardise on Tool Protocols

Replace custom API function-calling integrations with MCP-compatible interfaces where possible.

Step 3: Partition Static vs Dynamic State

Re-architect every system prompt you have into prefix and suffix components.

Step 4: Implement Context Quality Gates

Add governance to the context assembly pipeline before assembly reaches the model.

The Next Frontier

Naive RAG served a purpose. It made retrieval-augmented generation accessible and demonstrated that external knowledge could improve model accuracy on enterprise tasks. It was the starting point.

The starting point is not the destination.

Production enterprise AI infrastructure requires structured context, deterministic assembly, strict governance, and the economics that context caching provides. The teams that have crossed that line are running systems that perform at a different accuracy level, cost a fraction of what naive retrieval costs at scale, and can actually answer compliance and audit questions about what the model saw and why.

The teams still optimising chunk sizes are working on the wrong problem.

The era of dumping unstructured chunks into prompts is coming to an end.

The question that drives the next era of enterprise AI architecture is not “how do we retrieve more relevant chunks?” It is “how do we govern and structure what the model sees — so that every inference is accurate, auditable, cost-efficient, and deterministically repeatable?”

That question has a very different architecture as its answer.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.