Most of what an AI agent costs, in money and in waiting time, is input tokens it has already seen. A chat turn resends the same primer and answering rules every time. A Schedule Builder turn resends a firm's whole scheduling SOP on every model call, and one build makes hundreds of them. Prompt caching turns that repetition into a discount and a latency win, but only if the prompts and the model clients are built for it. We designed Connect, the agentic scheduling platform we built on top of Syncify, around that from the start.
This article covers how the caching works across Connect's agents: how prompts are ordered, how cache keys are chosen per role, how long sessions stay inside the context window, when cheap child models and lower reasoning effort take over, and how we check whether any of it is working. The chat-specific prompt assembly is walked through step by step in anatomy of an AI chat turn. This article is the cross-agent view.
How implicit prompt caching actually behaves
Connect's long-running agents (chat, the Info Sheet, the App Builder and the Schedule Builder) run on the OpenAI Agents SDK against Azure OpenAI through Azure AI Foundry, using the Responses API. The narrower agents run on a lighter runtime, but the same layout rules apply to their prompts. The provider offers implicit prompt caching. Three properties of it drive every decision below:
- It is a prefix cache. A request reuses cached computation up to the first byte that differs from an earlier request. One changed character near the top throws away everything after it.
- A cache key routes requests. Requests that share a key are steered to where the matching prefix is likely to be warm. The key is a routing hint. It is not a data boundary, and the cache still matches on content.
- Nothing is guaranteed. A hit is not promised, and how cached tokens are billed depends on the deployment. So a cache strategy that is never measured is a hope, not a strategy.
The practical consequence is that prompt layout becomes an architectural concern. Where a block sits in the prompt decides whether it is paid for once per session or once per call.
Static first, volatile last
Every Connect agent builds its prompt in two halves: an invariant prefix that is byte-identical across calls, then a volatile tail. The rule is simple: nothing that changes per call (timestamps, ids, retrieved passages, live state) may sit above anything that does not.
Here is an illustrative layout for the Schedule Builder, the agent with the longest sessions. Block sizes are orders of magnitude, not exact figures.
| Position | Block | Changes when | Typical size | Cache behaviour |
|---|---|---|---|---|
| 1 | Identity line with the project name | Never, within one build | One line | Prefix (per project) |
| 2 | Persona (the published version or the code default) | An admin publishes a new version | Small | Prefix |
| 3 | The firm's CPM scheduling SOP | Rarely, by deployment | Large | Prefix |
| 4 | The firm's machine-readable schedule standard | Rarely, by deployment | Large | Prefix |
| 5 | Build Spec language, gate discipline, working rules | By deployment | Medium | Prefix |
| 6 | Curated knowledge documents | An admin edits knowledge | Small to medium | Prefix |
| 7 | The approved Project Information Sheet | The sheet is re-approved | Medium | Prefix (per project) |
| 8 | Means & Methods, reference schedule, user setup | The user confirms new answers | Medium | Prefix, mostly stable |
| 9 | Saved schedule revision and open review comments | A revision saves, a comment opens | Small | Usually stable within a turn |
| 10 | Volatile state: the spec's shape, confirmed answers, last build report | Every repair pass | Small | Breaks the prefix here |
| 11 | Session history, tool calls and results | Every model call (append-only) | Grows | Cached within a pass |
| 12 | The new input | Every call | Small | Never cached |
Two details in that table are worth calling out.
The identity line names the project. That means the warm prefix is effectively per project, not shared across all builds. We accepted that. Within one build the prefix is byte-identical across a few dozen user turns and hundreds of model calls, which is where nearly all the repetition is.
The volatile state sits at the end of the instructions. The runtime appends a compact summary of the Build Spec and the last build report after the stable prompt on every pass. When that summary changes between repair passes, everything after it, including session history, is re-read. Within a single pass, though, the instructions are fixed and history only grows by appending, so each tool call in the loop hits on everything before the newest item. That is the shape you want: one cold read per pass, warm reads for every step inside it.
Chat follows the same pattern at a smaller scale: domain primer, curated knowledge (budgeted at 20,000 characters), the Syncify tools primer and the answering rules lead; retrieved passages, history and attachments trail. Toggles like "tools available" vary by the shape of a turn, not its content, so turns of the same shape share a prefix.
Split stable and volatile versions of the same thing
Curated knowledge and knowledge retrieved for the current message look like the same thing, but curated knowledge is stable for weeks and retrieved knowledge changes on every message. Carried together, they would put a volatile block high in the prompt and break the cache for everything below it. So they are separate prompt parameters in separate positions: curated knowledge in the prefix, retrieved knowledge in the tail. If two things share a label but not a rate of change, split them.
Tool definitions are part of the prefix
Tool schemas are sent with every request and count toward the cached prefix. After the Schedule Builder's first draft pass, the research and setup tools are dropped from the offer, because every strict tool schema offered slows the model's writing. That costs one cold read on the next pass. We took that trade on purpose: a smaller tool list speeds up every pass after it. The broader point is that changing the tool list mid-session is a cache decision, and it should be made knowingly.
Cache keys: one per role, never per tenant
Every model client in Connect stamps a prompt-cache key onto each request. The keys follow three rules: they carry a version, they carry no tenant or project data, and each distinct prompt role gets its own.
| Role | Key shape | Why |
|---|---|---|
| Chat parent | Versioned chat key, scoped by model | The follow-up generator runs on a smaller model with a tiny prompt; a shared slot would let it churn the main model's warm prefix |
| Schedule Builder lead | Fixed, versioned key | The lead resends the same prefix on every build call |
| Schedule Builder reviewers | One key per reviewer | Reviewers have their own prompts; sharing the lead's key would evict its warm prefix |
| Info Sheet orchestrator and normalizer | One fixed key each | Different prompts, different slots |
| Info Sheet section workers | Versioned prefix plus a short hash of project and section | Each section's history grows across calls; a stable hash routes it consistently without putting raw ids in the key |
The version suffix is the cheap insurance. When the shape of a prefix changes (a block moves, a new section is added), we bump it, and nothing can be routed toward a stale slot.
Keeping tenant data out of keys is not a security control, because the key is not a boundary. It keeps identifiers out of a field that ends up in provider logs and our own traces, and lets identical invariant prefixes share warmth across workspaces.
The model client: a thin subclass
All of this lives in small subclasses of the Agents SDK's Responses model, one per agent family, each overriding the method that builds the request. The pattern, written from scratch here:
class CachedModel extends ResponsesModel {
buildRequest(req, stream) {
const body = super.buildRequest(req, stream)
body.prompt_cache_key ??= KEY // a per-call override wins
body.prompt_cache_options = { mode: "implicit", ttl: "30m", ...body.prompt_cache_options }
return body
}
}
Two choices matter. Defaults are applied with "only if unset", so an agent that passes its own key through provider data (the reviewers, the section workers) is not overwritten. And the 30-minute implicit TTL fits how people work: long enough to cover a planner reading an agent's question and answering it, short enough not to hold prefixes nobody is using.
The Schedule Builder's client also sets a two-minute timeout on streamed requests that have not returned response headers, so a dead stream is abandoned and retried rather than held open for the SDK's much longer default. Retries are where a warm cache quietly pays off: the retried request resends the same prefix and reads most of it from cache.
Long sessions: output reserves and server-side compaction
Short chat turns run plain, with a per-turn output ceiling of 4,000 tokens. The long-running agents need more structure, so their clients set three things on every request: a fixed output reserve, truncation disabled, and server-side compaction above a threshold.
| Agent | Output reserve | Compaction threshold | Why the reserve is that size |
|---|---|---|---|
| Chat (sandboxed skills turns only) | 4,000 per turn | 100,000 | A safety valve for a turn that reads a large skill file |
| Project Info Sheet | 30,000 | 180,000 (in a 300,000 window) | Long, multi-section authoring turns |
| Schedule Builder | 16,000 | 240,000 | Spec edits are structured and compact |
| App Builder | 32,000 | 240,000 | A single staging call carries a complete HTML document |
Truncation is disabled on purpose. Silent truncation drops the oldest context without telling anyone, and the oldest context is often the instruction that matters. With truncation off and compaction on, the provider summarises older history explicitly once the threshold is crossed, and the run keeps going.
The reserve is part of the budget. A threshold must leave room for the reserve plus the next input. Setting the threshold close to the window and the reserve large is how a long tool loop ends in a context-length error on its last, most important call.
Compaction and caching interact. Compaction rewrites the history portion of the request, so the history after it reads cold on the next call. The instructions above it are untouched, which is another reason to keep the expensive, stable material in the instructions rather than in replayed messages.
Chat history is bounded twice
Chat does not replay a message array at all. Recent turns are folded into the volatile tail as a plain transcript, bounded to the last 10 turns and a 6,000-character budget filled newest-first. That keeps the volatile tail small, so the uncached part of each chat request stays cheap no matter how long the conversation runs. Attachments get their own 24,000-character budget for the same reason.
Cheap work goes to cheap models
Not every model call deserves the main model or deep reasoning. Connect pushes side tasks down:
| Task | Model | Reasoning effort | Other limits |
|---|---|---|---|
| Chat follow-up suggestions | Child model | Minimal | 200 output tokens, one turn |
| On-topic scope guard before an agent turn | Child model | Minimal | 200 output tokens, strict JSON verdict |
| Drawing reading for the Schedule Builder | Child model | Low by default | |
| App Builder critic | Child model | Low | Strict JSON output |
| Plain retrieval-and-answer chat turn | Main model | Low | 4,000 output tokens |
| Chat turn with tools | Main model | Configured default | 8 tool turns |
| Schedule Builder reviewers | Configured per deployment | Medium | Own cache key each |
| Info Sheet required sections | Configured per deployment | High | |
| Info Sheet advisory sections | Configured per deployment | Low | |
| Info Sheet normalizer | Configured per deployment | Medium | Own cache key |
A published agent version can override the thinking level, and when it does it wins. Otherwise effort is tiered by the shape of the work.
Stepping effort down for bulk repair
The Schedule Builder makes the most interesting effort decision. A build turn is a draft followed by repair passes, each fixing the failures the acceptance gates found. The runtime runs:
- The first draft at full configured effort, because a lower-effort draft writes a thinner schedule.
- Bulk repair passes, while more than five gate failures remain, one effort level lower.
- The last few failures at full effort again, because at low effort each last fix tends to knock another one loose.
The step-down has a floor: it never drops below low. It sheds latency on the passes where the work is mechanical and keeps the reasoning for the passes where a fix has to hold.
Measuring it: cache tokens in the ledger
Every agent run writes one row to agent_events, the Postgres trace ledger described in LLM agent tracing on a Postgres ledger. Cache usage gets first-class columns rather than being folded into input:
- Cache-read tokens, taken from the cached-token count in the response's input usage details.
- Cache-write tokens, recorded only when the provider reports them, and stored as unknown rather than zero when it does not.
- Reasoning tokens, kept separate from output for the same reason.
The executive dashboard sums those columns over any time range next to input tokens, cost and duration, so a cache hit rate is a division away: cache-read tokens over input tokens, per agent, per day. When a prompt change lands and that ratio drops, it shows up the next morning, not on the invoice.
The Info Sheet's client goes further, because its sessions are the most expensive. Each request carries hashed fingerprints of the full request, the instructions, the tool list, the settings and each input segment, plus the cache key, in request metadata. The trace processor stores them with each model turn. When a hit rate falls, comparing fingerprints between two consecutive calls shows exactly which part changed: a tool schema that reordered, a setting that flipped, an input segment that should have been stable. On endpoints that support it, the client also asks the provider to compare each request with the previous response and records its cache diagnostics.
Why this matters if you're building something similar
- Treat prompt order as architecture. Write the rule down (invariant first, volatile last, nothing per-call above the retrieved content) and review against it.
- Split blocks by rate of change, not by topic. Two kinds of "knowledge" that change at different speeds belong in different positions.
- Give every prompt role its own versioned key. Small side prompts sharing a slot with a large one will evict it. Bump the version when a prefix changes shape.
- Keep tenant data out of cache keys. The key is a routing hint that ends up in logs, not a boundary.
- Pair an output reserve with explicit compaction, and turn silent truncation off. Budget the threshold so the reserve and the next input always fit.
- Push side tasks to a child model at minimal effort, and step effort down for mechanical bulk work. Keep full effort for drafts and for the last fixes that have to hold.
- Log cache-read tokens per run. The provider does not guarantee hits. Your ledger is the only place you will find out whether the strategy works.
Where to go next
- How we built Connect: architecture of an enterprise AI platform, the series overview.
- Anatomy of an AI chat turn, the chat prompt assembled step by step.
- The AI Schedule Builder: build spec and gates, the agent with the longest sessions.
- LLM agent tracing on a Postgres ledger, where cache and reasoning tokens are recorded.
- Lessons from vector search on Cosmos DB for MongoDB vCore, the retrieval tier that feeds the volatile tail.
Running agents whose token bills are growing faster than their usage? Talk to us about it.
Frequently asked questions
What is prompt caching for AI agents?
It is a provider feature that reuses computation for the part of a request that matches an earlier one, byte for byte, from the start. Agents resend the same instructions on every model call, so a well-ordered prompt pays for its stable prefix once per session instead of once per call.
Should a prompt cache key include the tenant or customer id?
We keep tenant and project data out of cache keys. The key is a routing hint, not a data boundary, and it ends up in logs. Where a key needs to follow one growing history, Connect uses a short hash rather than raw ids.
Why disable truncation and use compaction instead?
Silent truncation drops the oldest context without telling anyone, and that is often the instruction that matters most. Server-side compaction summarises older history explicitly once a threshold is crossed, so a long tool loop keeps going without losing its footing.
How do you know whether prompt caching is actually working?
Every run records cache-read tokens from the provider's usage details in the trace ledger. Dividing cache-read tokens by input tokens per agent per day gives a hit rate, and request fingerprints show which part of a prompt changed when the rate drops.
When should an agent use lower reasoning effort?
For side tasks such as follow-up suggestions and scope checks, and for mechanical bulk work. Connect's Schedule Builder drafts at full effort, runs bulk repair passes one level lower, and returns to full effort for the last few fixes.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
Anatomy of an AI Chat Turn: Safety, Retrieval, Prompt Caching and Streaming, Step by Step
A stage-by-stage trace of a single Connect AI chat turn, showing what runs before the model call, what runs in parallel with it, and what must happen after it, with the latency and cost reasoning behind each step.
AI agents & MCP · October 9, 2026
Tracing LLM Agents in Postgres: Why We Built a Ledger Instead of OpenTelemetry
A deep dive on Connect's agent ledger: two recorders feeding one best-effort sink, per-turn rows with reasoning and cache tokens, a stricter table for MCP tool calls, and the admin views built on plain SQL.
Scheduling & P6 · October 9, 2026
Engineering an AI Schedule Builder: Build Specs, Gates and Repair Loops
The engineering behind the Connect Schedule Builder loop: the model edits a declarative Build Spec, code expands and gates it, reviewers and bounded repair passes drive it to done, and approval freezes verified evidence.