Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
How to Cut Your AI Agent’s Context Cost With GPT‑6 Prompt Caching
Latest   Machine Learning

How to Cut Your AI Agent’s Context Cost With GPT‑6 Prompt Caching

Last Updated on September 25, 2026 by Editorial Team

Author(s): Abhishek Badar

Originally published on Towards AI.

How to Cut Your AI Agent’s Context Cost With GPT‑6 Prompt Caching

If your AI agent repeatedly sends the same system instructions, documentation, tool definitions, coding standards, or repository context, you may be paying to process essentially the same tokens on every request.

GPT‑6 Sol and Luna launched on September 22 with API prices 50% below their GPT‑5.6 counterparts. More interesting for agent builders, GPT‑6 supports OpenAI’s newer prompt-caching controls, where cached reads cost 10% of normal input-token pricing.

That makes today a good time to restructure long agent prompts around a stable cached prefix.

The problem

Imagine an internal coding agent receives roughly this context on every request:

8,000 tokens — engineering rules
6,000 tokens — API documentation
4,000 tokens — tool definitions
2,000 tokens — current task

Only the last 2,000 tokens change frequently.

A common implementation nevertheless reconstructs everything:

const response = await openai.responses.create({
model: "gpt-6-sol",
input: [
{
role: "developer",
content: buildAgentInstructions()
},
{
role: "user",
content: currentTask
}
]
});

The model repeatedly sees roughly 20,000 input tokens.

Prompt caching changes the economics if you structure the request correctly.

OpenAI’s cache works on matching prompt prefixes. For GPT‑5.6 and later, an eligible prefix must contain at least 1,024 visible input tokens. A cache write costs 1.25× normal input pricing, but subsequent cached reads cost only 0.1×.

So reorganize the request as:

STABLE
────────────────────
agent instructions
coding conventions
API documentation
examples
shared reference data
────────────────────
CACHE BREAKPOINT
DYNAMIC
────────────────────
user
repository state
current task
recent tool results
────────────────────

The expensive context now has a reusable boundary.

Add an explicit cache breakpoint

GPT‑6 supports explicit cache breakpoints through the Responses API.

A simplified implementation looks like this:

import OpenAI from "openai";

const openai = new OpenAI();
const response = await openai.responses.create({
model: "gpt-6-sol",
prompt_cache_options: {
mode: "explicit"
},
input: [
{
role: "developer",
content: [
{
type: "input_text",
text: `
You are our repository engineering agent.
ENGINEERING RULES:
${engineeringRules}
API DOCUMENTATION:
${apiDocs}
ARCHITECTURE:
${architectureDocs}
EXAMPLES:
${examples}
`,
prompt_cache_breakpoint: {
mode: "explicit"
}
}
]
},
{
role: "user",
content: currentTask
}
]
});

The first request writes the stable prefix.

Later requests can reuse it:

Request 1
[████████████████████][task A]
↓
cache write

Request 2
[████████████████████][task B]
↑
cache read
Request 3
[████████████████████][task C]
↑
cache read

The critical detail is that the prefix must remain stable.

Don’t inject this into your cached instructions:

Current time: 10:43:21
User: 19382
Request ID: a8df...

You just destroyed prefix reuse.

Move changing information after the breakpoint.

OpenAI specifically recommends putting stable instructions, examples and reference material first and dynamic content afterward. Tool definitions and their ordering should also remain stable where possible.

The cost difference gets large quickly

Suppose your reusable context contains 20,000 input tokens.

GPT‑6 Sol currently costs $2 per million uncached input tokens.

Become a Medium member

Processing that prefix normally across ten requests costs roughly:

20,000 × 10 = 200,000 tokens
200,000 / 1,000,000 × $2≈ $0.40

With caching, the first write costs 1.25× and the next nine reads cost 0.1×.

Ignoring the changing suffix for simplicity:

20,000 × $2 × 1.25
─────────────────── ≈ $0.05
1,000,000

Nine cached reads:

180,000 × $2 × 0.1
────────────────── ≈ $0.036
1,000,000

Total:

without caching ≈ $0.40
with caching ≈ $0.086

That’s roughly a 78% reduction on that reusable portion of the input in this example.

OpenAI gives the same economics another way: across ten requests, one cache write plus nine complete cache reads costs 2.15× a single uncached processing of the prefix, versus 10× without caching.

The savings become much more consequential when an agent repeatedly carries tens of thousands of tokens of repository context, documentation and tool schemas.

Measure it instead of assuming it works

Caching fails silently from an architecture perspective: your application still works when you accidentally destroy cache reuse. You just pay more.

Log the cache metrics returned with each response, especially:

cached_tokens
cache_write_tokens

Then track:

cache hit rate =
cached input tokens
───────────────────
total eligible input tokens

If you’re running a long-context agent and consistently getting a low hit rate, inspect what changes between requests.

OpenAI now provides a prompt-cache diagnostics tool for supported GPT‑5.6-and-later models. It compares a current request with an earlier response and identifies differences in model settings, tools or input that prevented prefix reuse.

That gives you a useful production metric:

cost per completed agent task

rather than simply:

cost per model call

An agent may make 20 model calls to complete one coding task. Saving reusable context across those calls can matter more than shaving a few hundred tokens from individual prompts.

Try this today

Take one production workflow with a large prompt and inspect what actually changes between requests.

Move your stable instructions, examples, documentation and reference material to the beginning. Put user-specific state, timestamps, current tasks and changing tool results afterward.

Then add an explicit breakpoint and run the same workload 10–20 times.

Compare:

cached_tokens
input-token cost
p95 latency
cost per completed task

If your agent carries substantial repeated context, you should see the benefit quickly.

The mistake would be upgrading from GPT‑5.6 to GPT‑6 and stopping there.

The cheaper model helps once.

Restructuring the application so it stops repeatedly processing the same context helps on every request.

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.