AI agents & MCP · October 8, 2026 · 11 min read

Governing AI Agents in Construction: A Checklist for IT and Leadership

A practical checklist for putting AI agents into construction systems safely, covering credentials, least privilege, read-first rollout, tool routing, telemetry, evaluation and human approval, drawn from our own builds.

By Charley Forey, founder of Build Flows

Video walkthrough. Chapters and full transcript →

Governing AI agents in construction means deciding, before anything goes live, what each agent can see, what it can change, who approves the changes, and how you'll know what it did. In practice that comes down to eight controls: credentials stay on the server, least-privilege scopes, read-first rollout, tool routing, telemetry on every call, evaluation against rubrics and golden sets, human approval for writes and agent changes, and separate environments. This checklist covers each one, with the patterns we use in our own builds.

The audience is the people who get the call when an agent goes wrong: IT and security leads, controllers, project controls managers and the executives who signed off on the pilot. None of this needs a specific vendor. It does need decisions written down.

Why construction agents need their own governance

Agents in construction don't just answer questions. They call tools that touch Procore projects, Primavera P6 schedules, ERP cost data, CRM pipelines and client deliverables. One agent session can read a commitment, draft a change event and email a client in a few seconds.

Three things make construction riskier than most settings:

  • Many systems, many credentials. A typical contractor runs project management, scheduling, accounting, estimating and field tools, each with its own auth model and admin. Agents tend to pick up access to all of them.
  • Big API surfaces. Procore's public API alone produced 2,755 tools when we generated an MCP server from its OpenAPI specs. Hand an agent all of them and you get poor tool choices along with a much larger blast radius.
  • Money and contracts. Billing, retainage, change orders and baseline schedules have contractual weight. A wrong write isn't a typo. It can turn into a dispute.

Governance doesn't slow agents down. It's what lets IT approve putting them in front of real projects.

Key terms

TermWhat it means in practice
AgentA model that runs in a loop, choosing and calling tools to finish a task, not just producing text.
ToolOne callable operation an agent can use, such as "list RFIs" or "create change event," usually exposed through an MCP server.
MCP gatewayA single Model Context Protocol endpoint in front of many tool servers. It controls which tools each key can see, injects credentials and logs calls.
Least privilegeEvery agent, key and token gets the minimum access its named workflow needs, and nothing more.
Tool routingFinding and exposing only the tools that match the task's intent, instead of loading the whole catalog into context.
TelemetryA structured record of each tool call: who called it, through which key, against which server, with what input and output, and whether it succeeded.
RubricA written grading standard for agent output, used to score quality the same way every time.
Golden setA curated set of reference questions or tasks with known good answers. You re-run it whenever a prompt, model or tool changes.
Review queueA holding area where proposed changes to an agent (new skills, prompt edits, model swaps) wait for a human to approve them.

The governance checklist

Work through these roughly in order. The early steps make the later ones possible.

  1. Name the owner and the workflow. Every agent gets a named business owner and a written statement of what it's for: "Draft weekly schedule updates for the client," not "help with scheduling." If no one will own it, don't build it.
  2. Keep credentials on the server. API keys, OAuth refresh tokens and service principals never go into prompts, chat history or client-side config. The agent calls tools through a server, either an MCP gateway or an integration runtime, and that server injects auth at call time.
  3. Scope access to least privilege. Use project- or company-scoped tokens rather than org-wide admin credentials. Choose individual tools, not whole servers. A cost-reporting agent shouldn't be able to see "delete project."
  4. Start read-only. The first production release reads and reports. Add write scopes one named workflow at a time, each with owner sign-off.
  5. Route tools, don't dump them. For large APIs, expose a small set of discovery and call tools that find the right operation for the task, or a persona-based subset. Loading thousands of tool definitions on every model call wastes context and makes tool selection worse.
  6. Enforce policy at one chokepoint. Every call goes through the same path: validate the input, apply rate limits, check policy, guard idempotency, run the call, shape the output, redact sensitive fields and write an audit record. Rules that live in prompts get skipped. Rules in the dispatch path don't.
  7. Log every tool call. Record organization, user, key, server, tool, input, output, latency and outcome. That's how you debug a wrong answer and how you show IT what the agent actually did.
  8. Define quality before launch. Write the rubric, build the golden set and record a baseline score. Re-run the set on every change to prompts, models, tools or knowledge sources.
  9. Gate changes and writes behind human approval. Proposed agent changes go to a review queue. Write actions that carry financial or contractual weight need explicit confirmation or a validation loop.
  10. Separate environments. Build in sandbox or non-production first. Use offline artifacts, such as an exported XER file instead of a live P6 database, while you're still exploring.
  11. Plan the rollback. Version agent configurations so you can revert a prompt or model change. Know how to revoke a key in one step.
  12. Review usage regularly. Look at which agents, tools and models get used, by whom, at what token cost and with what success rate. Retire what nobody uses.

Credentials and least privilege

A common failure is an API key pasted into a prompt, a desktop config file or a low-code workflow. Once that happens you can't rotate the key without breaking things, you can't tell which agent used it, and anyone who can read the config can read the key.

The fix is architectural. In Tool Runtime, our MCP gateway, upstream authentication is stored with the saved server. An administrator then builds a composed MCP by picking servers and then individual tools from each one. That selection, plus rate limits, is bound to one API key. The agent's client gets an mcp.json that points at the gateway with that key. It never sees the upstream credentials, and it only sees the tools on its key.

Least privilege works the same way in agent platforms. In the Connect scheduling platform, workspaces carry role-based access (admin, member, guest), and feature flags switch capabilities on or off per workspace. That includes MCP access to the schedule builder agent and the choice between a core tool set and the full tool set. A client workspace gets only what that client needs.

Read-first rollout and human approval

Read-only agents carry most of the early value and very little risk. Reporting, search, summarization, schedule review and drafting a document for a person to send are all read-heavy. Write access should arrive one workflow at a time.

When an agent does write, add a human step that matches the stakes:

  • Draft, then approve. The agent writes a draft (a client update, a means-and-methods answer, a schedule revision) and a person publishes it. In Connect, client answers come back through a secure link and are marked client-verified, and every change to the means and methods keeps an audit trail.
  • Validation loops on creates. In our Procore MCP server, policies can block delete actions outright and require a checksum or validation loop before a create goes through.
  • Approve changes to the agent itself. Prompt edits, new skills, knowledge-source changes and model swaps are also changes to production. In Connect's executive admin portal they land in a review queue, and an admin approves them before they take effect.

A good agent will also decline to guess. When the information isn't in the source documents, the agent should ask a question or mark the field as unknown. It shouldn't invent a value.

Tool routing: governance through relevance

Routing usually gets treated as a performance concern. It's also a governance control. An agent can't misuse a tool it never sees.

Our Procore MCP server runs in three modes:

ModeWhat the agent seesWhen to use it
ExpandedAll 2,755 generated toolsClients that let you hand-pick tools per agent
HybridA subset chosen by persona, policy or request headers (for example, RFI-focused)Teams with a known, narrow job
RouterAbout 25 meta-tools for searching, describing, planning and callingGeneral agents under tight context limits

In router mode the agent calls tools like procore_search_tools and procore_call_tool. Retrieval combines keyword search (BM25) and embedding search, merged with reciprocal rank fusion, so the agent finds the right operation from what the user meant. Every call still goes through the same dispatch chokepoint: schema validation, rate limit, policy, idempotency, handler, output shaping, PII redaction, audit. Details are in routing 2,755 Procore API tools.

Multi-tenancy belongs here too. Request headers partition each organization's or team's instance so that no data crosses between tenants.

Telemetry and audit

If you can't answer "what did the agent do, with what data, on whose behalf?", you aren't running it in production yet. You're demoing it.

At minimum, capture:

  • Organization, user and API key (or workspace)
  • Server and tool called
  • Full input and output
  • Latency and success or failure
  • For agent loops: turns, tool sequence and the final answer

Tool Runtime logs the organization, user, API key, MCP server, input and output of every call, and its analytics show tool calls and success rates per API key. Connect keeps a log for every agent run, covering turns, tool calls, inputs and outputs, alongside usage analytics by agent, tool, user and model, including token cost. Those logs can be queried over MCP as well, so an engineer can ask an agent to find errors or improvement opportunities in other agents' runs.

Telemetry also feeds health monitoring. The Procore MCP server exposes health and metrics endpoints, so uptime, failing tools and unusual usage show up before a user reports them.

Evaluation: rubrics and golden sets

A model or prompt change can make one task better and quietly make another worse. Rubrics and golden sets catch that.

  • Rubric: write down what "good" means for each agent. For a project info sheet agent that might cover completeness against the drawings, citation of the source file, no invented values, and correct units and dates.
  • Golden set: collect real requests with reference answers your experts agree on. Re-run the set after every change and compare the scores.
  • Gate on the score. A change that lowers the score doesn't ship, however good it looked in a one-off test.

Connect grades its agents against rubrics and golden reference sets in the admin portal. A self-improvement agent can propose fixes (knowledge documents, prompt tweaks, model changes, grading rules), and those proposals go to the review queue for a person to approve. An autopilot mode can be switched on, but the default path is a person approving each change.

Domain checks count as evaluation too. Connect's schedule builder runs validation tests, produces a health score and keeps evidence for each activity showing how it reached each part of the schedule before you export the XER. For scheduling-specific checks such as logic gaps and critical path integrity, see AI for CPM scheduling and P6.

Environments and data posture

  • Build against sandbox or non-production tenants whenever the platform provides them. In our Procore MCP server, policies can limit writes to a sandbox.
  • Use offline artifacts to explore. Our P6 MCP server can analyze an XER file offline, so an agent can review a schedule without touching the live database.
  • Promote configuration the same way you promote code: versioned, reviewed and reversible.
  • Keep production keys separate from development keys, and rate-limit both.

Common mistakes

  1. Credentials in prompts or client configs. You can't rotate them, scope them or attribute their use.
  2. Exposing an entire API "to be safe." Larger tool lists lead to worse tool choices and a wider blast radius.
  3. Write access on day one. Get the read workflows trusted first.
  4. No owner. An agent with no named owner never gets reviewed, tuned or retired.
  5. Testing by vibes. A few good chats isn't evaluation. Without a golden set you can't tell whether a change helped.
  6. Logging only the final answer. You need the tool calls and their inputs and outputs to find where an answer went wrong.
  7. Policies in prompts. "Never delete anything" in a system prompt is a request. A policy check in the dispatch path is a control.
  8. Treating a public demo as diligence. Reference builds show the patterns. Your tenant, data and risk tolerance still need their own access plan.

How we approach it

Our governance page sets out our positions. Credentials stay server-side. Access starts read-only, and write scopes expand only for named workflows with owner sign-off. We route tools instead of dumping them, log every call, and build in non-production first. We won't ship an unattended agent that can write to production without an access policy, an owner and a rollback path.

Those positions come from our own builds:

  • Tool Runtime: one gateway, per-key tool composition, server-side auth, rate limits and per-call input/output logging.
  • Procore MCP: router and hybrid modes, personas and workflows, and a single dispatch chokepoint with policy, idempotency, PII redaction and audit.
  • Connect: agent configs with versioning, rubrics and golden sets, a review queue, usage analytics, and workspaces with roles and feature flags.

Each engagement ships with an access plan IT can review. It should answer the questions in this checklist: which systems, which scopes, which tools, who owns each agent, where the logs go and how to roll back.

Where to go next

Frequently asked questions

What does AI agent governance mean for a construction company?

It means deciding before anything goes live what each agent can see, what it can change, who approves the changes and how its actions are recorded. In practice that covers credential handling, least-privilege scopes, tool routing, telemetry, evaluation, human approval and separate environments.

Should AI agents have write access to Procore or our ERP?

Not at first. Start with read-only workflows such as reporting, search and drafting. Then add write scopes one named workflow at a time with owner sign-off, plus validation loops or human confirmation for actions with financial or contractual weight.

Where should API keys live when agents use tools?

On a server the agent calls, such as an MCP gateway or an integration runtime, which injects credentials at call time. Keys should never appear in prompts, chat history or client-side config files, because that makes them impossible to scope, rotate or attribute.

How do you test whether an AI agent is good enough for production?

Write a rubric that defines good output, build a golden set of real requests with expert-approved answers, and record a baseline score. Re-run the set after every prompt, model, tool or knowledge change, and do not ship changes that lower the score.

What should be logged for every AI agent tool call?

At minimum: organization, user, API key or workspace, the server and tool called, the full input and output, latency and success or failure. For agent loops, also record the turns and the order in which tools were called.

What is tool routing and why does it matter for security?

Tool routing exposes only the tools that match a task instead of a whole API catalog. It improves tool selection and shrinks the blast radius, because an agent cannot call an operation it was never given.

Next step

Have a problem like this?

Tell us the outcome you need. We'll tell you honestly how we'd approach it, and reply within two business days.

Keep learning