Build Flows

AI agents & MCP · October 9, 2026 · 10 min read

Layered AI Safety Guardrails: Injection Gates, Content Safety, Topic Classifiers and Structural Controls

How Connect screens every agent turn with four ordered checks, why the network checks fail open while other controls fail closed, how streamed output is corrected, and why read-only, workspace-pinned tools matter more than any prompt.

By Charley Forey, founder of Build Flows

Every AI product needs guardrails, and most teams reach for one of two extremes. Either a single moderation API call wrapped around the model, or a sprawling list of banned words that ends up refusing "what's the float on the critical path?" because "float" looked suspicious. Neither holds up in front of construction users, who talk about blasting, demolition and killing a schedule as part of an ordinary day.

In Connect, the agentic scheduling platform we built on top of Syncify, safety is a stack of layers, ordered cheapest and most certain first, with a deliberate decision about what each layer does when it fails. This article covers the runtime mechanics: the four input checks, output screening, daily caps, and the structural controls that make the prompt irrelevant as a security boundary. For the governance side (who approves what, rollout, audit and evaluation policy), see governing AI agents in construction. This is the code that runs on every turn.

One seam for every agent

Connect has several agents: the general chat, the Info Sheet agent, the Schedule Builder, Means & Methods, Updates, the Alert Drafter and the App Builder. All of them screen through one shared safety module. Each agent passes a small context (which agent, workspace, user, optionally the project and the assistant's previous message) and gets back a verdict: allow or decline, an internal reason code for telemetry, and a user-facing decline message.

Keeping it in one place means one set of tests, one set of thresholds and one place to change behavior. The internal reason (injection, abuse_input, a Content Safety category, an off-topic reason) goes to logs; the user only ever sees a polite, fixed message that points them back to their projects.

The input screen: four checks in cost order

OrderCheckCostCatchesOn failure
1Prompt-injection gateMicroseconds, no I/OLiteral attacks on the promptCannot fail (pure function)
2Abuse floorMicroseconds, no I/OUnambiguous slurs and directed harassmentCannot fail (pure function)
3Azure Content SafetyOne HTTP callHate, self-harm, sexual and violent content at or above a severity thresholdFails open
4Topic classifierOne small-model callRequests confidently unrelated to construction workFails open

The order is the point. The deterministic checks are free and certain, so they run first and short-circuit: an obvious injection attempt never costs a network call. The network checks are slower and probabilistic, so they run only for messages that passed the floor. The classifier, the most expensive and most subjective, runs last.

In sketch form:

screenInput(text):
  if matchesInjection(text): return decline("injection")
  if matchesAbuse(text):     return decline("abuse_input")
  category = contentSafety(text)       // null on any error
  if category: return decline("content_safety:" + category)
  if classifier enabled:
    verdict = classify(text, context)  // allow on any error
    if not verdict.allow: return decline("off_topic")
  return allow

1. The injection gate: narrow on purpose

The first check is a short list of patterns run against a normalized copy of the message. It targets well-known, copy-pasted phrasings that try to override or extract the system prompt.

It is deliberately allow-most. Every pattern has to be an unambiguous attack, not a suspicious word. "Ignore the weather delays, what's the finish date?" must pass, and the tests check exactly that. The gate does not try to catch off-topic or hostile questions. Widening it into a keyword topic filter would misfire on real domain questions ("where are we slipping?", "what's the float?"), and a false positive here is invisible: the user just gets refused and assumes the product does not work.

A pattern list is a cheap first filter, not the security boundary. It stops the copy-pasted attacks cheaply, and the structural controls described below make a successful injection far less valuable.

2. The abuse floor

The second check catches the least ambiguous categories of abuse: slurs and directed harassment. Anything subtler is left to Content Safety.

The floor exists so that some protection holds when Content Safety is unconfigured or down. It is intentionally compact. Mild job-site language ("these damn weather delays", "the crane plan is a nightmare") is not flagged, and the tests pin that. Growing this into a big dictionary would produce a brittle list that refuses real users and still misses things a trained classifier catches.

3. Azure Content Safety at a tuned threshold

The third check sends the message to Azure AI Content Safety for the Hate, SelfHarm, Sexual and Violence categories, using its graded severity output. A message is declined above a threshold we tuned for construction language.

The threshold is the interesting decision. At 2, ordinary construction risk talk trips the Violence category: demolition, blasting, "tear it down". At 4, real hate, sexual or self-harm content is still caught and site talk passes. If you adopt Content Safety, test the threshold against your own domain's vocabulary before shipping.

One practical detail: every call carries a timeout, so a slow moderation service cannot hold a turn open.

4. The topic classifier

The last check asks a smaller, cheaper model one question: is this plausibly about the user's construction work, or confidently unrelated (recipes, homework, poems, role-play)? It runs at minimal reasoning effort with a 200-token ceiling and a strict JSON schema output of allow plus reason, so there is nothing to parse loosely.

Two inputs make it accurate:

  • The assistant's last message (up to 600 characters). Without it, a terse reply like "yes, 12 floors" looks like a non-sequitur. With it, the classifier sees the user answering the agent's own question.
  • The project name, when there is one, for domain framing.

The instructions say "when in any doubt, allow", and only a confident "unrelated" blocks. If no AI is configured, or the classifier is switched off, this gate is skipped entirely and the turn relies on the first three layers plus the model's own instructions.

Why the network checks fail open

Both network checks fail open on errors, by design: a moderation outage counts as "nothing flagged" rather than as a refusal. That is enforced twice. Each client catches its own errors, and the safety module wraps each call again, so a collaborator that forgets to catch can never break a turn.

The reasoning is about which failure is worse for this product:

  • A false refusal is silent and corrosive. A superintendent who gets "I can't help with that" for a real question does not file a ticket. They stop using the tool.
  • A moderation outage should not take the assistant down. Content Safety and the classifier are external dependencies. Making them hard gates would couple Connect's availability to theirs.
  • The deterministic floor still runs. Failing open on the network layers degrades to the pure-function layers, not to nothing.
  • The blast radius is bounded elsewhere. The structural controls below limit what any message, screened or not, can make the agent do.

Fail-open is not a blanket policy, though. We chose a direction for every control:

ControlFailsWhy
Content Safety, topic classifierOpenFalse refusals are worse than a missed edge case; the floor still runs
Output screenOpenSame reasoning; the answer has already streamed
The screen as a whole (its promise rejects)ClosedAn unscreened turn must never reach the model
Agent paused by an administratorClosedA pause is a deliberate decision
Daily cap: count query errorsOpenIt is cost control, not a security boundary
Daily cap: count over the limitClosedA clean over-limit count is a hard refusal

The distinction in the third row matters. The individual network checks degrade gracefully inside the screen, but if the screen itself somehow errors, the turn fails rather than running unscreened.

Screening in parallel with retrieval

Moderation latency is real. In the chat agent, the input screen is started before retrieval and its result is awaited just before the first model call, so embedding, vector search and prompt assembly absorb most of its wait. The gate is preserved: no model call runs until the verdict is in, and a decline means the model is never called at all. The full ordering is in anatomy of an AI chat turn.

Output screening

After the model answers, the output screen runs the abuse floor and Content Safety over the full text. It catches the case where a model has been steered, perhaps by a poisoned document passage, into producing something it should not.

Streaming complicates this. Tokens already sent cannot be unsent. So for a streamed turn, a decline appends a correction frame to the stream, persists only the safe decline message (no citations), and drops the follow-up suggestions. For a non-streamed turn, the generated text is simply withheld. Follow-up generation runs concurrently with the output screen, so the screen adds little tail latency.

Daily message caps

Each user has a configurable daily cap on assistant messages, counted from UTC midnight. The policy (the limit and the refusal message) lives in the safety module; counting lives in each agent's own store, which keeps the module free of storage dependencies. The cap is checked in the turn preamble, before screening or any model work, so an over-limit user costs one count query.

Structural safety: the prompt is not a permission boundary

Screening reduces bad inputs. It cannot make an agent safe on its own, because a model can always be talked into trying something. The controls that matter most are the ones that make "trying" useless.

Read-only, workspace-pinned MCP proxies

Connect AI reads live Syncify data through a hosted MCP server, and the proxy between the agent and that server enforces two rules in code rather than in the prompt. It is read-only: it connects with the user's own token, keeps only tools that are both annotated and named as reads, and excludes new tools by default. And it is pinned to one workspace: any tool argument naming a workspace is overwritten with the mapped workspace id on every call, so the model cannot read across the tenant boundary by asking nicely. The full gate is described in the agent tool manifest article. For the general pattern, see MCP gateways, telemetry and the tool runtime.

Server-pinned scope everywhere else

The same idea runs through the rest of the agent surface:

  • Native tools have the workspace and user pinned from the turn's trusted scope. The model supplies only in-workspace selectors like a project or schedule id, and every call re-authorizes against the user's current project access.
  • Tools are filtered by role before the model sees them. A tool that is not in the list cannot be talked into running.
  • Retrieval scope is computed on the server from the user's visible projects. The model never chooses what it can search.
  • Untrusted text is fenced. Retrieved passages, attachments and history are wrapped and labelled as data, never instructions.

These controls are why the injection gate can afford to be narrow. A successful jailbreak still runs with the asking user's permissions, read-only, inside one workspace. More on the authorization side in multi-tenant RBAC with a single authorization chokepoint.

Testing the layers

The safety module ships with focused unit tests that use stub moderators and classifiers which count their calls. They lock the properties that matter: injection and abuse decline with zero network calls made; a Content Safety flag or off-topic verdict declines; a clean construction question passes; the screen fails open when Content Safety or the classifier throws; the topic gate is skipped when no classifier is configured; output screening declines poisoned text and fails open on error; the daily cap trips exactly at the limit; and the Content Safety client maps severities correctly and fails open on bad responses.

The domain cases matter as much as the attack cases. Tests assert that "what's the float on the critical path?" and "these damn weather delays are killing the schedule" pass. A guardrail suite that only tests what gets blocked will drift toward blocking too much.

Why this matters if you're building something similar

  • Order checks by cost and certainty. Pure functions first, network calls only for what survives.
  • Keep deterministic filters narrow. Their job is unambiguous cases; trained classifiers handle nuance.
  • Tune moderation thresholds to your domain. Construction vocabulary trips default thresholds.
  • Give classifiers conversational context. The previous assistant message turns false positives into passes.
  • Choose a failure direction for every control, and write it down.
  • Plan for streamed output. Decide in advance how you correct an answer the user has already seen.
  • Put real security in structure, not prompts. Read-only filters, server-pinned scope and role-filtered tools limit damage regardless of what gets through.
  • Test false positives as hard as attacks.

Where to go next

Want guardrails designed for your own agents and data? Talk to us about your build.

Frequently asked questions

Why do the Content Safety and classifier checks fail open?

A false refusal silently breaks a real question, and a moderation outage should not take the assistant down. The deterministic checks still run, and structural controls bound what any message can do.

Why not block on a list of suspicious keywords?

Construction language is full of words that look alarming out of context, like float, blast, demolition or killing the schedule. A keyword filter misfires on real questions, so the deterministic gates only target unambiguous attacks and slurs.

What Content Safety threshold does Connect use?

Severity 4 on the four-level scale, across Hate, SelfHarm, Sexual and Violence. That is high enough to let ordinary construction risk talk through and low enough to catch real abuse.

How do you moderate an answer that has already streamed?

The output screen runs on the full answer. If it declines, a correction is appended to the stream, only the safe decline message is stored, and follow-up suggestions are dropped.

How do you stop an agent reaching another tenant's data through MCP?

The proxy connects as the user, keeps only read-only tools, and overwrites any workspace argument with the mapped workspace id on every call, so the model cannot choose a different workspace.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning