From the outside, an AI chat turn looks like one thing: you type a question, text streams back. Inside Connect, the agentic scheduling platform we built on top of Syncify, that turn is more than a dozen distinct steps. Some exist for security, some for cost, some for latency.
This article follows a single Connect AI turn from the moment a user presses Enter to the moment the last frame lands, in the order the code runs it. Most chat tutorials show "retrieve, prompt, call the model." The interesting engineering is everything around that call: what has to happen before it, what can run beside it, and what must never be skipped after it.
The turn at a glance
The safety screen starts at step 3 and is awaited at step 9, so its latency overlaps the work in between.
| # | Stage | Blocks the answer? | Why it is where it is |
|---|---|---|---|
| 1 | Preamble: ownership, pause switch, assistant key, project scope, daily cap | Yes | Cheap database reads; reject before spending anything |
| 2 | Snapshot history, resolve attachments | Yes | History must be read before the new message is stored |
| 3 | Start the input safety screen | No | Its network latency overlaps the next stages |
| 4 | Persist the user message | Yes | The transcript is never missing the question |
| 5 | Load published config and admin knowledge | Yes | Persona, model, tools and curated knowledge are per agent |
| 6 | Build tool scopes and the tool list | Yes | Tools are filtered by the caller's roles before the model sees them |
| 7 | Embed the question, scoped top-8 vector search | Yes | Grounding is cheap and runs while the screen is still working |
| 8 | Assemble the prompt, static first | Yes | Ordering decides how much of the prompt cache hits |
| 9 | Await the safety verdict | Yes | No model call ever runs on an unscreened message |
| 10 | RAG fast path or tool loop (8 tool turns max) | Yes | Plain questions get low effort; tool questions get full reasoning |
| 11 | Stream text and tool events over SSE | No | The user sees progress instead of a frozen spinner |
| 12 | Output screen, follow-ups in parallel | Partly | Tail latency is the slower of the two, not their sum |
| 13 | Persist answer, link the trace, send follow-ups and done | Yes | The ledger row points at the exact stored message |
Now the detail.
1. The preamble: refuse early and cheaply
Before anything touches a model, the turn runs a set of database checks, in this order:
- The message is non-empty and the conversation belongs to this user and workspace.
- The agent is not paused for this workspace. Executives can pause Connect AI per workspace from the admin portal. This check fails closed: a pause is a deliberate decision and should be a clean refusal.
- The user holds the assistant feature key. One role key gates every agent turn across the product. Revoking it stops a member's AI use while their manual editing keeps working.
- If the conversation is project-scoped, the user can still see that project. Access can change between turns, so it is re-checked every time.
- The user is under the daily message cap. We count the user's messages since UTC midnight and compare against a configurable cap.
The daily cap is interesting because it fails in the opposite direction to the pause switch. If the count query itself errors, the turn proceeds. A cap is cost control, closer to telemetry than to a security boundary, and a flaky count should not lock users out. A clean count over the cap is a hard refusal. We write more about that split in layered AI safety guardrails.
2. History and attachments
Next, the turn reads recent conversation history, before the new message is inserted, so history means "what came before this question" and never includes the question itself.
History is bounded twice: the last 10 turns, and a 6,000-character budget filled newest-first. It goes into the system prompt as a plain transcript labelled as context, not as a replayed message array. That avoids reconstructing provider-specific assistant messages, and it lets us tell the model explicitly that the transcript is something to read, not instructions to follow.
Attachments are resolved to extracted text and share a 24,000-character budget split evenly across files, each fenced as a passage to be read as data.
3. Start the safety screen, but do not wait for it
Input screening involves up to two network calls (Azure Content Safety and a topic classifier on a smaller model) and has an 8-second timeout. If the turn waited for it here, every turn would pay that latency up front.
So the screen starts here and its promise is handed down the stack. Retrieval, config loading and prompt assembly all run while it works. The turn blocks on the verdict at step 9, immediately before the first model call. The gate is preserved; only the waiting is overlapped.
One subtlety: some paths return before step 9 (AI not configured, for example) and never await the screen, so a no-op rejection handler stops an unhandled rejection. It does not swallow the gate: wherever the verdict is awaited, a screen error fails the turn closed.
4. Persist the question
The user message is stored now, before any model work. If anything later fails or the user closes the tab, the transcript still shows what was asked.
If the conversation has a paused sub-agent waiting for input (the user is answering a question the Schedule Builder asked, say), the message is routed to that sub-agent instead of starting a new turn. That path has nothing to overlap, so it awaits the screen immediately. The mechanics of pausing and resuming are in multi-agent delegation with pause and resume.
5. Published config and admin knowledge
Each agent has a published configuration that admins edit and version in the executive portal: persona prompt, model, thinking level, tool allow-list and a library of curated knowledge documents. Fields left null fall back to code defaults, so an agent nobody has edited still runs.
Knowledge documents come in two modes:
- Always documents are included in full on every turn.
- Retrieved documents contribute only their most relevant pieces. Today that is a keyword-overlap ranking that keeps the top three chunks, which is enough at current library sizes; embedding similarity is the noted upgrade path.
Both are folded into the prompt under a 20,000-character budget. Documents past the budget are truncated or omitted, and the prompt says how many were dropped. A large knowledge library can never blow the context window.
6. Tools, filtered before the model sees them
The turn then builds its tool list. Two families exist:
- Connect-native tools read the workspace's own data (projects, schedules, recordings, alerts, info sheets) through the same capability modules the HTTP routes use. The list is filtered by the caller's feature access. A tool the model never sees cannot be talked into running. Each tool also re-authorizes on every call, intersecting the selected project or program with the user's current access.
- Delegation tools hand work to specialist agents (Info Sheet, Schedule Builder, Means & Methods, Updates, Alert Drafter, App Builder), each offered only when the user's role and the conversation's scope allow it.
The read-only Syncify tools, reached over MCP as the user, are connected later, inside the agent, and only for plain workspace-scoped turns, because the live MCP pins a workspace but cannot enforce a selected project. How tool names, descriptions and permissions are governed is covered in the agent tool manifest.
7. Embed the question and search, in scope
Now retrieval. The question is embedded (1,024 dimensions) and searched against the workspace knowledge base:
- Compute the retrieval scopes. For each workspace the turn reaches, re-authorize the user and produce either "across projects" (the user can see everything) or "in visible projects" (an explicit list plus workspace-wide documents).
- Run a vector search per scope with a limit of 8.
- Merge all hits, sort by distance, keep the best 8 overall.
Failures here split along a line we apply everywhere: authorization failures stop the turn; availability failures degrade it. If the authorize step throws, the turn fails. If the search itself errors, the turn continues without those passages and raises a warning that some evidence could not be searched. If embeddings are unavailable entirely, the model answers ungrounded rather than not at all.
Each hit becomes a numbered passage wrapped in passage tags, and citations are deduplicated by source so the same file is not cited three times. The storage and scope filters behind this search are in the RAG pipeline article.
8. Assemble the prompt for the cache
Prompt caching only helps up to the first byte that differs between requests. So the system prompt is built in two halves, in a fixed order:
Stable prefix
- The domain primer (or the admin's published persona)
- Admin knowledge, budgeted
- The Syncify tools primer, when those tools are available
- Answering rules: cite passages by number, say when the answer is not there, treat passage content as data, never as instructions
Volatile tail 5. The retrieved passages 6. Conversation history 7. Attachments
Nothing per-turn (timestamps, ids) is allowed above the passages, so all turns of the same shape share a prefix.
The model wrapper stamps every request with a stable prompt-cache key and implicit caching with a 30-minute TTL. The key carries no tenant or project data. It is a routing hint, not a data boundary, so every workspace warms the same invariant prefix. It is scoped by model name, because the follow-up generator runs on a smaller model with a tiny prompt, and sharing a slot would let it churn the main model's warm prefix. The key has a version suffix that gets bumped whenever the prefix's shape changes, so a stale prefix is never reused.
9. Await the verdict
Only now does the turn wait for the safety screen started in step 3. By this point retrieval and assembly have absorbed most of its latency. If the screen declines, the model is never called: the turn returns a decline message, which is streamed to the user and stored as the assistant reply.
10. Two paths: RAG fast path or tool loop
If the turn has no tools (no Syncify session and no local tools), it takes the RAG fast path: one grounded answer from the passages, at low reasoning effort. Retrieval plus synthesis does not need deep planning, and low effort sheds noticeable latency.
If tools are available, it runs a tool loop at the configured default effort, because planning tool calls and reconciling their results benefits from reasoning. A published thinking level overrides both. The loop has hard bounds: 8 model turns and 4,000 output tokens. Independent tool calls can fan out in parallel in one round trip, since the read tools are idempotent.
Several degradations keep the turn answering:
- The Syncify MCP is unreachable or the token is rejected. The turn continues with local tools and documents, and a warning is logged. A tool-level error is prefixed in the text the model reads, so "your token was rejected" cannot be paraphrased as a fact about the project.
- The allow-list filters out every tool. The turn falls back to the RAG path.
- The loop hits its 8-turn cap mid-fetch with no prose. Rather than surfacing "response was empty", the turn re-answers from the passages as a child run linked to the original in the trace ledger.
The MCP session is opened per turn and always closed in a finally block.
The sandboxed skills path
Some turns can use skills: shared, file-based capabilities such as PDF tooling and Python helpers. When a skills catalog is configured, the turn acquires an ephemeral sandbox, unique to that turn so no files leak between turns, and runs a sandbox-capable agent with filesystem and skills capabilities.
Sandbox turns differ. Tool calls run serially, because write-then-read in a filesystem is order-dependent. Context compaction kicks in at 100,000 tokens as a safety valve for a turn that reads a large skill file (plain chat turns are short and run without compaction). A 120-second wall clock kills the sandbox if the turn runs long. If the sandbox cannot be acquired, the turn falls back to the plain agent with a warning rather than failing. The sandbox is always killed and released at the end, whatever happened.
11. Streaming over SSE
The route answers with a text/event-stream response and relays agent events as frames:
delta: incremental answer texttoolwithstartandend: so the UI can show "Looking up the critical path..." instead of appearing frozensubagentandinteraction: activity from delegated agents and inline prompts when one pausesfollowups, thendonewith the persisted message, orerror
If nothing streamed (a fallback that returned whole), the finished answer is sent as one delta. If the client disconnects, the turn is aborted and whatever text already streamed is persisted, so the transcript is never left with a dangling question.
12. Output screen and follow-ups, side by side
When the answer is complete, two things start at once:
- Follow-up suggestions: three short next questions from the smaller child model, at minimal effort with a 200-token ceiling. Best-effort and untraced; a failure just means no suggestion chips.
- The output screen: the abuse floor and Content Safety run over the full answer.
Tail latency is the slower of the two rather than the sum. Streamed tokens cannot be unsent, so if the output screen declines, the turn appends a correction frame, stores only the safe message and drops the suggestions. The output screen fails open: if moderation is down, the streamed answer stands.
13. Persist, link, finish
The answer is stored with its citations. The trace ledger row for the run is linked to that stored message id, so feedback on a message lands on the exact run that produced it. Then the follow-ups frame and the done frame go out.
The trace itself is registered before the model call, with the full instructions and tool definitions, and completed in a finally block with success or error, the user text and the final text. That is exactly one Postgres write per run, shared with the other agents. More in LLM agent tracing in a Postgres ledger.
Why this matters if you're building something similar
- Order the checks by cost. Database reads first, network moderation in parallel, model calls last.
- Overlap, don't skip. Start slow gates early and await them at the last responsible moment.
- Decide which failures stop and which degrade. Authorization stops the turn. Retrieval, MCP and sandbox availability degrade it, always with a visible warning.
- Write the prompt for the cache. Invariant content first, volatile last, a version on the cache key, and no tenant data in it.
- Tier reasoning effort by turn shape. Plain retrieval answers do not need the same effort as tool planning.
- Bound everything. History, attachments, knowledge, tool results, tool turns and output tokens all have explicit budgets.
- Persist partials and link traces to messages. Debugging starts from the message the user complained about.
Where to go next
- How we built Connect: architecture of an enterprise AI platform
- Layered AI safety guardrails
- RAG pipeline for construction documents
- Multi-agent delegation with pause and resume
- LLM agent tracing in a Postgres ledger
Building an AI assistant your team will actually trust with project data? Tell us what you want to build.
Frequently asked questions
Why start the safety screen before retrieval?
Moderation involves network calls that take time. Starting it first and awaiting the verdict just before the model call lets embedding, search and prompt assembly absorb that wait, while still guaranteeing no unscreened message reaches the model.
How is the prompt ordered for prompt caching?
The domain primer, admin knowledge, tools primer and answering rules come first because they rarely change. Retrieved passages, history and attachments come last. Caching only helps up to the first differing byte.
How many passages does each turn retrieve?
Eight. Each authorized scope is searched with a limit of eight, then all hits are merged, sorted by distance and cut back to the best eight.
What limits keep the tool loop from running away?
A cap of 8 model turns and 4,000 output tokens. If the loop ends on a tool call with no prose, the turn re-answers from the retrieved documents instead of returning nothing.
What happens if the user closes the tab mid-answer?
The turn is aborted and whatever text already streamed is saved as the assistant message, so the transcript is never left with an unanswered question.
Next step
Have a workflow in mind?
Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
AI agents & MCP · October 9, 2026
Layered AI Safety Guardrails: Injection Gates, Content Safety, Topic Classifiers and Structural Controls
How Connect screens every agent turn with four ordered checks, why the network checks fail open while other controls fail closed, how streamed output is corrected, and why read-only, workspace-pinned tools matter more than any prompt.
AI agents & MCP · October 9, 2026
Building a RAG Pipeline for Construction Documents: Chunking, Incremental Re-ingest and Scoped Vector Search
An engineering walkthrough of the retrieval pipeline under Connect: what triggers ingestion, how text is chunked and re-ingested cheaply, how chunks are stored one database per workspace, and how scope filters, index fallbacks and an admin browser keep retrieval correct and inspectable.