Build Flows

AI agents & MCP · October 9, 2026 · 10 min read

The Agent Tool Manifest: Declaring, Gating and Testing Every Tool an AI Agent Can Call

A look at Connect's tool manifest: one registry pinned by tests, writes as proposals, capability-to-feature RBAC checked at offer and call time, per-agent allow-lists narrowed by published configs, and a read-only proxy for a remote MCP server.

By Charley Forey, founder of Build Flows

Every tool you give an AI agent is a second door into your application. If the HTTP route behind "list schedules" checks roles but the agent tool that reads the same data doesn't, a member without schedule access can just ask the agent. Your access model is then optional. The same goes for writes. A tool that changes tenant data has to be reviewed with the same care as a new API endpoint, and in most codebases nothing forces that review.

In Connect, the AI platform we built on top of Syncify, we handled this the way we handle HTTP routes. There is one declared registry of every tool an agent may call, and a test suite fails the build when the real tool surface drifts from it. This article walks through that registry, the metadata each tool carries, how per-agent allow-lists are derived and then narrowed, and how we proxy an external MCP server we don't control without handing it any write power.

One registry, enforced by a test

The registry is a single typed array, TOOL_MANIFEST. Each row describes one verb:

FieldWhat it meansWho reads it
nameStable snake_case id. It's also the telemetry key in the agent ledgerThe model, the ledger, the admin panel
descriptionPlain English, written onceThe LLM tool schema, the admin "Tools" panel, generated docs
mutatesTrue only if the tool directly changes tenant stateA safety test, and the reviewer
agentsWhich agents may call it: chat, pis, app_builderDerived allow-lists
capabilityWhich capability module owns the implementationThe RBAC mapping and the layering test

Writing the description once has a practical benefit. The text the model sees, the text an executive sees in the admin tools view, and the docs can never disagree. Every agent's default allow-list is derived from the agents field. Nobody keeps it by hand.

Our planning doc also listed a cost field (cheap, moderate, expensive) for budgeting. The shipped type leaves it out, because nothing reads it yet. A field nobody reads goes stale and starts to mislead. When rate limiting needs it, it goes in along with the code that uses it.

The same rule applies to rows. A tool gets a row only when its adapter ships. A planned tool with no implementation would make the admin panel advertise something the agent can't do.

What's in it today

The manifest declares about sixty tools. Two more groups sit beside it: the Schedule Builder's own runtime tools, and the read-only tools discovered at runtime from the Syncify MCP. Roughly:

CategoryAgentCount (order of magnitude)Mutating
Info Sheet investigation and writespisabout a dozen3, each validated by the Info Sheet module's state rules
Chat reads (projects, schedules, Updates, recordings, alerts, playbooks, connector records, live Procore GETs)chatabout 200
Chat delegations to other agentschat62, both drafts a person must review
Shared app data tools (declared queries, propose_query)chat + app_builder30
App Builder draft, preview and propose loopapp_builderabout 151 (a sandboxed preview, never an install)
Schedule Builder runtime toolsschedule_builderabout a dozen, declared separatelySpec edits inside its own build session
Syncify MCP proxy toolschatvaries with the remote server0, enforced

Six of the manifest's sixty or so tools are flagged mutates. That ratio is deliberate.

Writes are proposals

The core rule: a new state-changing tool should stage a reviewable proposal, not perform the change. The Schedule Builder proposes a revision and the user accepts it. The App Builder stages a draft, publishes it as an uninstalled preview version, and then calls propose_app, which pauses for the user to install or discard. A child critic reviews the draft first and blocks a broken one. propose_query lets a builder ask for a new declared SQL query. The query is checked mechanically (it must filter on the workspace and project parameters) and stored as pending. It can't run until a super-admin approves it.

None of those count as direct mutations, and the manifest says so. The test suite pins the exact list of tools that do mutate, and a comment explains each one:

  • Three Info Sheet writes: record topic dispositions, save the profile, refresh the profile. Each one goes through the Info Sheet module's own state machine.
  • publish_draft_preview: it publishes an immutable app version that the user can preview live. That bundle runs in the app sandbox with no network path and no credentials, and nothing is installed until the user accepts. It replaced two older tools that published and installed in one unreviewed step.
  • run_means_methods: the Means & Methods drafter has no dry run, so it writes. But every answer it writes is marked as an assumption or a client question, never as confirmed, and an expert verifies it.
  • generate_project_update: it creates a client Update in a generating state. A client only sees it after a person publishes it.

If you add a seventh, the test fails, and someone has to edit that list with a reason. That small piece of friction is the whole point.

Delegation tools are an interesting case. run_schedule_builder is marked mutates: false even though a schedule can change at the end of it. The flag follows who commits the write. The builder stages a proposal, and the write happens only when the user's accept goes back to that builder. The writes happen in the child's turn, through the child's own declared tools.

Capability is a permission, not a label

The capability field does more than document ownership. It maps to a feature key in the RBAC model (projects, schedules, updates, recordings, apps, and so on). Chat offers a tool only if the caller can reach that feature. We check this twice:

  1. When the tool list is built for the turn. A member with no schedules access never sees list_schedules in the schema.
  2. When the tool is called. Before every call, the adapter re-authorizes: agent access, whether the agent is paused for the workspace, current features, and the allowed project set. If a role changes mid-turn, later calls are revoked even though the model already saw the tool.

Reaching a feature at normal (view) level is enough. The agent should be exactly as capable as the caller's own UI, no more and no less.

Some capabilities map to no feature, which leaves their tools ungated. That set is the dangerous one, because a typo ("schedule" for "schedules") would quietly fall into it. So the test pins the set to exactly two entries: agent-admin (playbooks are how the agent knows its job, not tenant data) and alerts (open to every workspace member at the route layer too, so the tool matches the UI). Another test builds the chat tools for a caller with no features and checks that no tenant-data tool is in the list.

Inside each call, the adapter also checks the objects the model names. A projectId must be in both the conversation's selected scope and the user's allowed projects. A scheduleId is looked up and its project checked the same way. A schedule comparison that names explicit snapshot ids must name snapshots of schedules the user can see. Workspace-wide tools (alert rules, propose_query) are refused outright when the conversation is narrowed to one project. Every result is capped at 20,000 characters before it reaches the model.

Allow-lists narrowed by published configs

The manifest sets the ceiling for each agent. An executive can lower it per agent from the admin portal. A published agent config has an enabledTools column:

  • NULL means use the defaults, which is every tool the manifest grants that agent.
  • A list means only these. Unknown names are rejected when the config is saved, so a typo can't publish an agent that silently has no tools.
  • An empty list means none for the App Builder. Chat treats an empty list like NULL. The meaning of "empty" is a per-agent decision, so it has to be written down and tested, not assumed.

The difference between NULL and empty sounds pedantic, but the UI makes it easy to get wrong. A "Use default tools" checkbox and a "0 tools selected" state look alike. A Playwright test drives the App Builder's settings page at phone and desktop widths, in light and dark themes. It saves an empty selection and checks the draft carries an empty array. It ticks "Use default tools" and checks the draft carries null. Then it clears the selection, picks one tool, and checks the draft carries exactly that one. The validator also maps a renamed legacy tool to its new name, so older published configs keep working.

Chat's allow-list governs only Connect-defined tools. The Syncify proxy tools are discovered per user at runtime, so they have no stable names to list. They keep their own read-only gate, described next.

The Schedule Builder declares its tools in its own runtime

The Schedule Builder registers its own tools inside its runtime: reading drawings with vision, recording answers to the build spec, applying the firm's templates, editing the spec, asking the user. These tools close over a live build session, so they aren't manifest rows. They're still declared, as a literal list next to the other agents' declared tools. A test reads the runtime's source, pulls out the tools it registers in order, and checks that they match the declared list exactly. This matters because a published config's allow-list filters the builder's tools by name. A stale name in the declared list would quietly leave a published builder without that tool.

Proxying an MCP server you don't control

Chat can also read live data from the hosted Syncify MCP server. We don't control that server, and its tool list can change without us knowing. The proxy treats every listed tool as untrusted and decides for itself:

  1. It connects as this user. The bearer token is the user's own Syncify token, so every call is limited by that user's own permissions on the remote side.
  2. Prefix filter. Only tools under the server's read namespace are considered. Anything new outside it is excluded by default.
  3. Annotation check. A tool that declares readOnlyHint false, or destructiveHint true, is refused.
  4. Name check. A tool whose name contains a write verb (create, update, delete, set, write, patch, put, post, archive, assign, move, rename, import, upload, sync) is refused, even if it claims to be read-only. The MCP spec itself calls annotations hints from a server you may not trust.
  5. Workspace pin. The user's token can see every workspace they belong to, which is wider than the Connect workspace they're chatting in. If a tool's schema has a workspace argument, the proxy overwrites it on every call with the workspace mapped to this Connect workspace. Last write wins, so whatever the model asked for is replaced.
  6. Bounded, labelled results. Each result is cut to 20,000 characters. A result flagged isError gets a "tool error:" prefix, so "no projects match" can't be confused with "your token was rejected".
  7. The accepted surface is reported. Refused tool names, and the case where nothing usable was offered, go to the turn's warnings. A toolset that quietly shrinks looks exactly like a model that stopped calling tools, so this is the only place that drift shows up.

In a minimal form, the gate is a handful of lines:

isReadOnly(tool):
  if not tool.name.startsWith(READ_PREFIX): return false
  if tool.annotations.readOnlyHint == false: return false
  if tool.annotations.destructiveHint == true: return false
  return not WRITE_VERB.matches(tool.name)

Two rules are enforced here and not in the prompt, because a prompt is not a permission boundary. The proxy also holds an open HTTP session, so the caller has to close it when the turn ends. That's in the contract too.

Why this matters if you're building something similar

  • Treat tools like routes. Keep one declared registry and a test that pins the real surface to it. Review comes free when a tool can't ship without a row.
  • Make mutation a flagged exception. Default to propose-and-approve, then pin the short list of direct writers in a test with a reason for each.
  • Tie each tool to an existing permission. Map tools to the same feature keys your UI uses, check at offer time and again at call time, and pin the ungated set so a typo can't grow it.
  • Keep scope out of the model's hands. Overwrite workspace and tenant arguments from the session. Never trust them from tool input.
  • Distrust remote tool metadata. Combine annotations with a name heuristic, default new tools to excluded, and report what you refused.
  • Leave out fields nothing reads yet. A smaller schema that's true beats a larger one that's partly made up.

Where to go next

Want an agent tool surface your security team can review line by line? Plan your build.

Frequently asked questions

What metadata should every agent tool declare?

At minimum a stable name, one description used everywhere, whether it directly mutates tenant state, which agents may call it, and which capability owns it. Those fields drive allow-lists, RBAC and review.

Why are most write tools proposals rather than direct mutations?

A proposal stages a change a human accepts, so a model mistake costs a rejected card rather than a bad write. The few direct writers are pinned in a test with a written reason for each.

How do you stop an agent reaching data the user cannot see?

Map each tool to the same feature key the UI uses, offer it only when the caller holds that feature, re-authorize on every call, and check every project or schedule id the model supplies against the user's allowed set.

Can you safely proxy a third-party MCP server's tools?

Yes, if you treat its metadata as untrusted: accept only a known prefix, refuse tools annotated as writes or named like writes, overwrite tenant arguments from the session, cap result size, and log what you refused.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning