Build Flows

AI agents & MCP · October 10, 2026 · 12 min read

Prompt Caching for AI Agents: How We Structure Prompts and Model Clients in Connect

A cross-agent look at how Connect orders prompts, chooses cache keys, keeps long sessions inside the context window, routes side tasks to cheap models and measures cache hits in its trace ledger.

By Charley Forey, founder of Build Flows

Most of what an AI agent costs, in money and in waiting time, is input tokens it has already seen. A chat turn resends the same primer and answering rules every time. A Schedule Builder turn resends a firm's whole scheduling SOP on every model call, and one build makes hundreds of them. Prompt caching turns that repetition into a discount and a latency win, but only if the prompts and the model clients are built for it. We designed Connect, the agentic scheduling platform we built on top of Syncify, around that from the start.

This article covers how the caching works across Connect's agents: how prompts are ordered, how cache keys are chosen per role, how long sessions stay inside the context window, when cheap child models and lower reasoning effort take over, and how we check whether any of it is working. The chat-specific prompt assembly is walked through step by step in anatomy of an AI chat turn. This article is the cross-agent view.

How implicit prompt caching actually behaves

Connect's long-running agents (chat, the Info Sheet, the App Builder and the Schedule Builder) run on the OpenAI Agents SDK against Azure OpenAI through Azure AI Foundry, using the Responses API. The narrower agents run on a lighter runtime, but the same layout rules apply to their prompts. The provider offers implicit prompt caching. Three properties of it drive every decision below:

  1. It is a prefix cache. A request reuses cached computation up to the first byte that differs from an earlier request. One changed character near the top throws away everything after it.
  2. A cache key routes requests. Requests that share a key are steered to where the matching prefix is likely to be warm. The key is a routing hint. It is not a data boundary, and the cache still matches on content.
  3. Nothing is guaranteed. A hit is not promised, and how cached tokens are billed depends on the deployment. So a cache strategy that is never measured is a hope, not a strategy.

The practical consequence is that prompt layout becomes an architectural concern. Where a block sits in the prompt decides whether it is paid for once per session or once per call.

Static first, volatile last

Every Connect agent builds its prompt in two halves: an invariant prefix that is byte-identical across calls, then a volatile tail. The rule is simple: nothing that changes per call (timestamps, ids, retrieved passages, live state) may sit above anything that does not.

Here is an illustrative layout for the Schedule Builder, the agent with the longest sessions. Block sizes are orders of magnitude, not exact figures.

PositionBlockChanges whenTypical sizeCache behaviour
1Identity line with the project nameNever, within one buildOne linePrefix (per project)
2Persona (the published version or the code default)An admin publishes a new versionSmallPrefix
3The firm's CPM scheduling SOPRarely, by deploymentLargePrefix
4The firm's machine-readable schedule standardRarely, by deploymentLargePrefix
5Build Spec language, gate discipline, working rulesBy deploymentMediumPrefix
6Curated knowledge documentsAn admin edits knowledgeSmall to mediumPrefix
7The approved Project Information SheetThe sheet is re-approvedMediumPrefix (per project)
8Means & Methods, reference schedule, user setupThe user confirms new answersMediumPrefix, mostly stable
9Saved schedule revision and open review commentsA revision saves, a comment opensSmallUsually stable within a turn
10Volatile state: the spec's shape, confirmed answers, last build reportEvery repair passSmallBreaks the prefix here
11Session history, tool calls and resultsEvery model call (append-only)GrowsCached within a pass
12The new inputEvery callSmallNever cached

Two details in that table are worth calling out.

The identity line names the project. That means the warm prefix is effectively per project, not shared across all builds. We accepted that. Within one build the prefix is byte-identical across a few dozen user turns and hundreds of model calls, which is where nearly all the repetition is.

The volatile state sits at the end of the instructions. The runtime appends a compact summary of the Build Spec and the last build report after the stable prompt on every pass. When that summary changes between repair passes, everything after it, including session history, is re-read. Within a single pass, though, the instructions are fixed and history only grows by appending, so each tool call in the loop hits on everything before the newest item. That is the shape you want: one cold read per pass, warm reads for every step inside it.

Chat follows the same pattern at a smaller scale: domain primer, curated knowledge (budgeted at 20,000 characters), the Syncify tools primer and the answering rules lead; retrieved passages, history and attachments trail. Toggles like "tools available" vary by the shape of a turn, not its content, so turns of the same shape share a prefix.

Split stable and volatile versions of the same thing

Curated knowledge and knowledge retrieved for the current message look like the same thing, but curated knowledge is stable for weeks and retrieved knowledge changes on every message. Carried together, they would put a volatile block high in the prompt and break the cache for everything below it. So they are separate prompt parameters in separate positions: curated knowledge in the prefix, retrieved knowledge in the tail. If two things share a label but not a rate of change, split them.

Tool definitions are part of the prefix

Tool schemas are sent with every request and count toward the cached prefix. After the Schedule Builder's first draft pass, the research and setup tools are dropped from the offer, because every strict tool schema offered slows the model's writing. That costs one cold read on the next pass. We took that trade on purpose: a smaller tool list speeds up every pass after it. The broader point is that changing the tool list mid-session is a cache decision, and it should be made knowingly.

Cache keys: one per role, never per tenant

Every model client in Connect stamps a prompt-cache key onto each request. The keys follow three rules: they carry a version, they carry no tenant or project data, and each distinct prompt role gets its own.

RoleKey shapeWhy
Chat parentVersioned chat key, scoped by modelThe follow-up generator runs on a smaller model with a tiny prompt; a shared slot would let it churn the main model's warm prefix
Schedule Builder leadFixed, versioned keyThe lead resends the same prefix on every build call
Schedule Builder reviewersOne key per reviewerReviewers have their own prompts; sharing the lead's key would evict its warm prefix
Info Sheet orchestrator and normalizerOne fixed key eachDifferent prompts, different slots
Info Sheet section workersVersioned prefix plus a short hash of project and sectionEach section's history grows across calls; a stable hash routes it consistently without putting raw ids in the key

The version suffix is the cheap insurance. When the shape of a prefix changes (a block moves, a new section is added), we bump it, and nothing can be routed toward a stale slot.

Keeping tenant data out of keys is not a security control, because the key is not a boundary. It keeps identifiers out of a field that ends up in provider logs and our own traces, and lets identical invariant prefixes share warmth across workspaces.

The model client: a thin subclass

All of this lives in small subclasses of the Agents SDK's Responses model, one per agent family, each overriding the method that builds the request. The pattern, written from scratch here:

class CachedModel extends ResponsesModel {
  buildRequest(req, stream) {
    const body = super.buildRequest(req, stream)
    body.prompt_cache_key ??= KEY            // a per-call override wins
    body.prompt_cache_options = { mode: "implicit", ttl: "30m", ...body.prompt_cache_options }
    return body
  }
}

Two choices matter. Defaults are applied with "only if unset", so an agent that passes its own key through provider data (the reviewers, the section workers) is not overwritten. And the 30-minute implicit TTL fits how people work: long enough to cover a planner reading an agent's question and answering it, short enough not to hold prefixes nobody is using.

The Schedule Builder's client also sets a two-minute timeout on streamed requests that have not returned response headers, so a dead stream is abandoned and retried rather than held open for the SDK's much longer default. Retries are where a warm cache quietly pays off: the retried request resends the same prefix and reads most of it from cache.

Long sessions: output reserves and server-side compaction

Short chat turns run plain, with a per-turn output ceiling of 4,000 tokens. The long-running agents need more structure, so their clients set three things on every request: a fixed output reserve, truncation disabled, and server-side compaction above a threshold.

AgentOutput reserveCompaction thresholdWhy the reserve is that size
Chat (sandboxed skills turns only)4,000 per turn100,000A safety valve for a turn that reads a large skill file
Project Info Sheet30,000180,000 (in a 300,000 window)Long, multi-section authoring turns
Schedule Builder16,000240,000Spec edits are structured and compact
App Builder32,000240,000A single staging call carries a complete HTML document

Truncation is disabled on purpose. Silent truncation drops the oldest context without telling anyone, and the oldest context is often the instruction that matters. With truncation off and compaction on, the provider summarises older history explicitly once the threshold is crossed, and the run keeps going.

The reserve is part of the budget. A threshold must leave room for the reserve plus the next input. Setting the threshold close to the window and the reserve large is how a long tool loop ends in a context-length error on its last, most important call.

Compaction and caching interact. Compaction rewrites the history portion of the request, so the history after it reads cold on the next call. The instructions above it are untouched, which is another reason to keep the expensive, stable material in the instructions rather than in replayed messages.

Chat history is bounded twice

Chat does not replay a message array at all. Recent turns are folded into the volatile tail as a plain transcript, bounded to the last 10 turns and a 6,000-character budget filled newest-first. That keeps the volatile tail small, so the uncached part of each chat request stays cheap no matter how long the conversation runs. Attachments get their own 24,000-character budget for the same reason.

Cheap work goes to cheap models

Not every model call deserves the main model or deep reasoning. Connect pushes side tasks down:

TaskModelReasoning effortOther limits
Chat follow-up suggestionsChild modelMinimal200 output tokens, one turn
On-topic scope guard before an agent turnChild modelMinimal200 output tokens, strict JSON verdict
Drawing reading for the Schedule BuilderChild modelLow by default
App Builder criticChild modelLowStrict JSON output
Plain retrieval-and-answer chat turnMain modelLow4,000 output tokens
Chat turn with toolsMain modelConfigured default8 tool turns
Schedule Builder reviewersConfigured per deploymentMediumOwn cache key each
Info Sheet required sectionsConfigured per deploymentHigh
Info Sheet advisory sectionsConfigured per deploymentLow
Info Sheet normalizerConfigured per deploymentMediumOwn cache key

A published agent version can override the thinking level, and when it does it wins. Otherwise effort is tiered by the shape of the work.

Stepping effort down for bulk repair

The Schedule Builder makes the most interesting effort decision. A build turn is a draft followed by repair passes, each fixing the failures the acceptance gates found. The runtime runs:

  1. The first draft at full configured effort, because a lower-effort draft writes a thinner schedule.
  2. Bulk repair passes, while more than five gate failures remain, one effort level lower.
  3. The last few failures at full effort again, because at low effort each last fix tends to knock another one loose.

The step-down has a floor: it never drops below low. It sheds latency on the passes where the work is mechanical and keeps the reasoning for the passes where a fix has to hold.

Measuring it: cache tokens in the ledger

Every agent run writes one row to agent_events, the Postgres trace ledger described in LLM agent tracing on a Postgres ledger. Cache usage gets first-class columns rather than being folded into input:

  • Cache-read tokens, taken from the cached-token count in the response's input usage details.
  • Cache-write tokens, recorded only when the provider reports them, and stored as unknown rather than zero when it does not.
  • Reasoning tokens, kept separate from output for the same reason.

The executive dashboard sums those columns over any time range next to input tokens, cost and duration, so a cache hit rate is a division away: cache-read tokens over input tokens, per agent, per day. When a prompt change lands and that ratio drops, it shows up the next morning, not on the invoice.

The Info Sheet's client goes further, because its sessions are the most expensive. Each request carries hashed fingerprints of the full request, the instructions, the tool list, the settings and each input segment, plus the cache key, in request metadata. The trace processor stores them with each model turn. When a hit rate falls, comparing fingerprints between two consecutive calls shows exactly which part changed: a tool schema that reordered, a setting that flipped, an input segment that should have been stable. On endpoints that support it, the client also asks the provider to compare each request with the previous response and records its cache diagnostics.

Why this matters if you're building something similar

  • Treat prompt order as architecture. Write the rule down (invariant first, volatile last, nothing per-call above the retrieved content) and review against it.
  • Split blocks by rate of change, not by topic. Two kinds of "knowledge" that change at different speeds belong in different positions.
  • Give every prompt role its own versioned key. Small side prompts sharing a slot with a large one will evict it. Bump the version when a prefix changes shape.
  • Keep tenant data out of cache keys. The key is a routing hint that ends up in logs, not a boundary.
  • Pair an output reserve with explicit compaction, and turn silent truncation off. Budget the threshold so the reserve and the next input always fit.
  • Push side tasks to a child model at minimal effort, and step effort down for mechanical bulk work. Keep full effort for drafts and for the last fixes that have to hold.
  • Log cache-read tokens per run. The provider does not guarantee hits. Your ledger is the only place you will find out whether the strategy works.

Where to go next

Running agents whose token bills are growing faster than their usage? Talk to us about it.

Frequently asked questions

What is prompt caching for AI agents?

It is a provider feature that reuses computation for the part of a request that matches an earlier one, byte for byte, from the start. Agents resend the same instructions on every model call, so a well-ordered prompt pays for its stable prefix once per session instead of once per call.

Should a prompt cache key include the tenant or customer id?

We keep tenant and project data out of cache keys. The key is a routing hint, not a data boundary, and it ends up in logs. Where a key needs to follow one growing history, Connect uses a short hash rather than raw ids.

Why disable truncation and use compaction instead?

Silent truncation drops the oldest context without telling anyone, and that is often the instruction that matters most. Server-side compaction summarises older history explicitly once a threshold is crossed, so a long tool loop keeps going without losing its footing.

How do you know whether prompt caching is actually working?

Every run records cache-read tokens from the provider's usage details in the trace ledger. Dividing cache-read tokens by input tokens per agent per day gives a hit rate, and request fingerprints show which part of a prompt changed when the rate drops.

When should an agent use lower reasoning effort?

For side tasks such as follow-up suggestions and scope checks, and for mechanical bulk work. Connect's Schedule Builder drafts at full effort, runs bulk repair passes one level lower, and returns to full effort for the last few fixes.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning