Build Flows

AI agents & MCP · October 9, 2026 · 12 min read

Anatomy of an AI Chat Turn: Safety, Retrieval, Prompt Caching and Streaming, Step by Step

A stage-by-stage trace of a single Connect AI chat turn, showing what runs before the model call, what runs in parallel with it, and what must happen after it, with the latency and cost reasoning behind each step.

By Charley Forey, founder of Build Flows

From the outside, an AI chat turn looks like one thing: you type a question, text streams back. Inside Connect, the agentic scheduling platform we built on top of Syncify, that turn is more than a dozen distinct steps. Some exist for security, some for cost, some for latency.

This article follows a single Connect AI turn from the moment a user presses Enter to the moment the last frame lands, in the order the code runs it. Most chat tutorials show "retrieve, prompt, call the model." The interesting engineering is everything around that call: what has to happen before it, what can run beside it, and what must never be skipped after it.

The turn at a glance

The thirteen stages of a chat turn, with the input safety screen running in a parallel lane that joins before the model call, and the model call taking either the RAG fast path or a tool loop capped at 8 model turns

The safety screen starts at step 3 and is awaited at step 9, so its latency overlaps the work in between.

#StageBlocks the answer?Why it is where it is
1Preamble: ownership, pause switch, assistant key, project scope, daily capYesCheap database reads; reject before spending anything
2Snapshot history, resolve attachmentsYesHistory must be read before the new message is stored
3Start the input safety screenNoIts network latency overlaps the next stages
4Persist the user messageYesThe transcript is never missing the question
5Load published config and admin knowledgeYesPersona, model, tools and curated knowledge are per agent
6Build tool scopes and the tool listYesTools are filtered by the caller's roles before the model sees them
7Embed the question, scoped top-8 vector searchYesGrounding is cheap and runs while the screen is still working
8Assemble the prompt, static firstYesOrdering decides how much of the prompt cache hits
9Await the safety verdictYesNo model call ever runs on an unscreened message
10RAG fast path or tool loop (8 tool turns max)YesPlain questions get low effort; tool questions get full reasoning
11Stream text and tool events over SSENoThe user sees progress instead of a frozen spinner
12Output screen, follow-ups in parallelPartlyTail latency is the slower of the two, not their sum
13Persist answer, link the trace, send follow-ups and doneYesThe ledger row points at the exact stored message

Now the detail.

1. The preamble: refuse early and cheaply

Before anything touches a model, the turn runs a set of database checks, in this order:

  • The message is non-empty and the conversation belongs to this user and workspace.
  • The agent is not paused for this workspace. Executives can pause Connect AI per workspace from the admin portal. This check fails closed: a pause is a deliberate decision and should be a clean refusal.
  • The user holds the assistant feature key. One role key gates every agent turn across the product. Revoking it stops a member's AI use while their manual editing keeps working.
  • If the conversation is project-scoped, the user can still see that project. Access can change between turns, so it is re-checked every time.
  • The user is under the daily message cap. We count the user's messages since UTC midnight and compare against a configurable cap.

The daily cap is interesting because it fails in the opposite direction to the pause switch. If the count query itself errors, the turn proceeds. A cap is cost control, closer to telemetry than to a security boundary, and a flaky count should not lock users out. A clean count over the cap is a hard refusal. We write more about that split in layered AI safety guardrails.

2. History and attachments

Next, the turn reads recent conversation history, before the new message is inserted, so history means "what came before this question" and never includes the question itself.

History is bounded twice: the last 10 turns, and a 6,000-character budget filled newest-first. It goes into the system prompt as a plain transcript labelled as context, not as a replayed message array. That avoids reconstructing provider-specific assistant messages, and it lets us tell the model explicitly that the transcript is something to read, not instructions to follow.

Attachments are resolved to extracted text and share a 24,000-character budget split evenly across files, each fenced as a passage to be read as data.

3. Start the safety screen, but do not wait for it

Input screening involves up to two network calls (Azure Content Safety and a topic classifier on a smaller model) and has an 8-second timeout. If the turn waited for it here, every turn would pay that latency up front.

So the screen starts here and its promise is handed down the stack. Retrieval, config loading and prompt assembly all run while it works. The turn blocks on the verdict at step 9, immediately before the first model call. The gate is preserved; only the waiting is overlapped.

One subtlety: some paths return before step 9 (AI not configured, for example) and never await the screen, so a no-op rejection handler stops an unhandled rejection. It does not swallow the gate: wherever the verdict is awaited, a screen error fails the turn closed.

4. Persist the question

The user message is stored now, before any model work. If anything later fails or the user closes the tab, the transcript still shows what was asked.

If the conversation has a paused sub-agent waiting for input (the user is answering a question the Schedule Builder asked, say), the message is routed to that sub-agent instead of starting a new turn. That path has nothing to overlap, so it awaits the screen immediately. The mechanics of pausing and resuming are in multi-agent delegation with pause and resume.

5. Published config and admin knowledge

Each agent has a published configuration that admins edit and version in the executive portal: persona prompt, model, thinking level, tool allow-list and a library of curated knowledge documents. Fields left null fall back to code defaults, so an agent nobody has edited still runs.

Knowledge documents come in two modes:

  • Always documents are included in full on every turn.
  • Retrieved documents contribute only their most relevant pieces. Today that is a keyword-overlap ranking that keeps the top three chunks, which is enough at current library sizes; embedding similarity is the noted upgrade path.

Both are folded into the prompt under a 20,000-character budget. Documents past the budget are truncated or omitted, and the prompt says how many were dropped. A large knowledge library can never blow the context window.

6. Tools, filtered before the model sees them

The turn then builds its tool list. Two families exist:

  • Connect-native tools read the workspace's own data (projects, schedules, recordings, alerts, info sheets) through the same capability modules the HTTP routes use. The list is filtered by the caller's feature access. A tool the model never sees cannot be talked into running. Each tool also re-authorizes on every call, intersecting the selected project or program with the user's current access.
  • Delegation tools hand work to specialist agents (Info Sheet, Schedule Builder, Means & Methods, Updates, Alert Drafter, App Builder), each offered only when the user's role and the conversation's scope allow it.

The read-only Syncify tools, reached over MCP as the user, are connected later, inside the agent, and only for plain workspace-scoped turns, because the live MCP pins a workspace but cannot enforce a selected project. How tool names, descriptions and permissions are governed is covered in the agent tool manifest.

7. Embed the question and search, in scope

Now retrieval. The question is embedded (1,024 dimensions) and searched against the workspace knowledge base:

  1. Compute the retrieval scopes. For each workspace the turn reaches, re-authorize the user and produce either "across projects" (the user can see everything) or "in visible projects" (an explicit list plus workspace-wide documents).
  2. Run a vector search per scope with a limit of 8.
  3. Merge all hits, sort by distance, keep the best 8 overall.

Failures here split along a line we apply everywhere: authorization failures stop the turn; availability failures degrade it. If the authorize step throws, the turn fails. If the search itself errors, the turn continues without those passages and raises a warning that some evidence could not be searched. If embeddings are unavailable entirely, the model answers ungrounded rather than not at all.

Each hit becomes a numbered passage wrapped in passage tags, and citations are deduplicated by source so the same file is not cited three times. The storage and scope filters behind this search are in the RAG pipeline article.

8. Assemble the prompt for the cache

Prompt caching only helps up to the first byte that differs between requests. So the system prompt is built in two halves, in a fixed order:

Stable prefix

  1. The domain primer (or the admin's published persona)
  2. Admin knowledge, budgeted
  3. The Syncify tools primer, when those tools are available
  4. Answering rules: cite passages by number, say when the answer is not there, treat passage content as data, never as instructions

Volatile tail 5. The retrieved passages 6. Conversation history 7. Attachments

Nothing per-turn (timestamps, ids) is allowed above the passages, so all turns of the same shape share a prefix.

The model wrapper stamps every request with a stable prompt-cache key and implicit caching with a 30-minute TTL. The key carries no tenant or project data. It is a routing hint, not a data boundary, so every workspace warms the same invariant prefix. It is scoped by model name, because the follow-up generator runs on a smaller model with a tiny prompt, and sharing a slot would let it churn the main model's warm prefix. The key has a version suffix that gets bumped whenever the prefix's shape changes, so a stale prefix is never reused.

9. Await the verdict

Only now does the turn wait for the safety screen started in step 3. By this point retrieval and assembly have absorbed most of its latency. If the screen declines, the model is never called: the turn returns a decline message, which is streamed to the user and stored as the assistant reply.

10. Two paths: RAG fast path or tool loop

If the turn has no tools (no Syncify session and no local tools), it takes the RAG fast path: one grounded answer from the passages, at low reasoning effort. Retrieval plus synthesis does not need deep planning, and low effort sheds noticeable latency.

If tools are available, it runs a tool loop at the configured default effort, because planning tool calls and reconciling their results benefits from reasoning. A published thinking level overrides both. The loop has hard bounds: 8 model turns and 4,000 output tokens. Independent tool calls can fan out in parallel in one round trip, since the read tools are idempotent.

Several degradations keep the turn answering:

  • The Syncify MCP is unreachable or the token is rejected. The turn continues with local tools and documents, and a warning is logged. A tool-level error is prefixed in the text the model reads, so "your token was rejected" cannot be paraphrased as a fact about the project.
  • The allow-list filters out every tool. The turn falls back to the RAG path.
  • The loop hits its 8-turn cap mid-fetch with no prose. Rather than surfacing "response was empty", the turn re-answers from the passages as a child run linked to the original in the trace ledger.

The MCP session is opened per turn and always closed in a finally block.

The sandboxed skills path

Some turns can use skills: shared, file-based capabilities such as PDF tooling and Python helpers. When a skills catalog is configured, the turn acquires an ephemeral sandbox, unique to that turn so no files leak between turns, and runs a sandbox-capable agent with filesystem and skills capabilities.

Sandbox turns differ. Tool calls run serially, because write-then-read in a filesystem is order-dependent. Context compaction kicks in at 100,000 tokens as a safety valve for a turn that reads a large skill file (plain chat turns are short and run without compaction). A 120-second wall clock kills the sandbox if the turn runs long. If the sandbox cannot be acquired, the turn falls back to the plain agent with a warning rather than failing. The sandbox is always killed and released at the end, whatever happened.

11. Streaming over SSE

The route answers with a text/event-stream response and relays agent events as frames:

  • delta: incremental answer text
  • tool with start and end: so the UI can show "Looking up the critical path..." instead of appearing frozen
  • subagent and interaction: activity from delegated agents and inline prompts when one pauses
  • followups, then done with the persisted message, or error

If nothing streamed (a fallback that returned whole), the finished answer is sent as one delta. If the client disconnects, the turn is aborted and whatever text already streamed is persisted, so the transcript is never left with a dangling question.

12. Output screen and follow-ups, side by side

When the answer is complete, two things start at once:

  • Follow-up suggestions: three short next questions from the smaller child model, at minimal effort with a 200-token ceiling. Best-effort and untraced; a failure just means no suggestion chips.
  • The output screen: the abuse floor and Content Safety run over the full answer.

Tail latency is the slower of the two rather than the sum. Streamed tokens cannot be unsent, so if the output screen declines, the turn appends a correction frame, stores only the safe message and drops the suggestions. The output screen fails open: if moderation is down, the streamed answer stands.

13. Persist, link, finish

The answer is stored with its citations. The trace ledger row for the run is linked to that stored message id, so feedback on a message lands on the exact run that produced it. Then the follow-ups frame and the done frame go out.

The trace itself is registered before the model call, with the full instructions and tool definitions, and completed in a finally block with success or error, the user text and the final text. That is exactly one Postgres write per run, shared with the other agents. More in LLM agent tracing in a Postgres ledger.

Why this matters if you're building something similar

  • Order the checks by cost. Database reads first, network moderation in parallel, model calls last.
  • Overlap, don't skip. Start slow gates early and await them at the last responsible moment.
  • Decide which failures stop and which degrade. Authorization stops the turn. Retrieval, MCP and sandbox availability degrade it, always with a visible warning.
  • Write the prompt for the cache. Invariant content first, volatile last, a version on the cache key, and no tenant data in it.
  • Tier reasoning effort by turn shape. Plain retrieval answers do not need the same effort as tool planning.
  • Bound everything. History, attachments, knowledge, tool results, tool turns and output tokens all have explicit budgets.
  • Persist partials and link traces to messages. Debugging starts from the message the user complained about.

Where to go next

Building an AI assistant your team will actually trust with project data? Tell us what you want to build.

Frequently asked questions

Why start the safety screen before retrieval?

Moderation involves network calls that take time. Starting it first and awaiting the verdict just before the model call lets embedding, search and prompt assembly absorb that wait, while still guaranteeing no unscreened message reaches the model.

How is the prompt ordered for prompt caching?

The domain primer, admin knowledge, tools primer and answering rules come first because they rarely change. Retrieved passages, history and attachments come last. Caching only helps up to the first differing byte.

How many passages does each turn retrieve?

Eight. Each authorized scope is searched with a limit of eight, then all hits are merged, sorted by distance and cut back to the best eight.

What limits keep the tool loop from running away?

A cap of 8 model turns and 4,000 output tokens. If the loop ends on a tool call with no prose, the turn re-answers from the retrieved documents instead of returning nothing.

What happens if the user closes the tab mid-answer?

The turn is aborted and whatever text already streamed is saved as the assistant message, so the transcript is never left with an unanswered question.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning