AI agents & MCP · May 10, 2026 · 8 min read

Routing 2,755 Procore API Tools Through MCP Without Drowning the Model

A walkthrough of the Procore MCP server architecture: generated tools for every endpoint, router-mode meta-tools, hybrid retrieval, and a single dispatch chokepoint for safety and audit.

By Charley Forey, founder of Build Flows

Video walkthrough. Chapters and full transcript →

Procore exposes a very large REST API. When we generated an MCP server from Procore's combined OpenAPI specification, it produced 2,755 tools: one for every endpoint, across RFIs, submittals, budgets, commitments, inspections, directory, documents and the rest of the platform.

Having every endpoint available is the easy part. The hard part is letting an AI agent use that surface without loading 2,755 tool definitions into every request, and without handing a language model unrestricted write access to a live construction project. This article walks through how our open-source Procore MCP server handles that: discovery, routing, a single governed dispatch path, and the observability that keeps it running in production.

The problem: context overload at Procore scale

An MCP server tells the client what tools exist by sending each tool's name, description and input schema. The client passes those definitions to the model as part of the prompt, and it does that on every model call.

At a few dozen tools that's fine. At 2,755 it breaks down:

  • Token cost. The tool catalog alone can take up most of the context window before the user has asked anything.
  • Selection quality. When a model has thousands of near-identical options (list RFIs, list RFI replies, list project RFIs, list RFI assignees), it's more likely to pick the wrong one or invent parameters.
  • Access posture. A flat dump means every agent can see and call every write and delete endpoint. No IT or security team should sign off on that.

Most MCP servers today map one tool to one API endpoint and return the whole list to the agent. As Charley puts it in the walkthrough, that's an early and somewhat naive way to build them. It works for small APIs. It doesn't hold up for an enterprise platform with thousands of operations.

The approach: routing over dumping

"Routing over dumping" is one of our design principles, and this server is a direct example of it. You don't show the model the whole API. You give it a small set of meta-tools that let it search for the right operation, read that operation's details, and then call it. The full catalog stays on the server, and the model only sees what the current task needs.

The server supports three modes, chosen per request in the MCP authentication headers:

ModeWhat the agent seesWhen to use it
ExpandedAll 2,755 generated toolsClients that let you hand-pick which tools reach the model
HybridA subset filtered by persona, policy or focus area (for example, RFIs only)Role-specific agents with a known scope
RouterAbout 25 meta-tools for discovery, description, calling and planningGeneral-purpose agents working across Procore

In router mode, the agent works through tools such as procore_search_tools, procore_call_tool and procore_plan_dag. They fall into a few groups:

  1. Discovery and description. Search the tool catalog by intent, then pull the full schema and context for the few candidates that match.
  2. Calling. Execute the selected operation, including chunking and paginating large calls and running calls in sequence.
  3. Resolution. When the agent doesn't have enough information, it asks the user a clarifying question instead of guessing.
  4. Multi-step planning. Work out which tools a task needs, in what order, and what data each step needs from the step before it.

The planning group is where most of the useful behavior comes from. Real Procore work is rarely one call. Answering an RFI question might mean resolving the company, then the project, then the RFI, then its responses and attachments. Each step depends on IDs returned by the one before it.

How discovery finds the right tool

Search quality decides whether router mode works. If procore_search_tools returns the wrong candidates, the agent fails no matter how good the model is.

We use hybrid retrieval over vectorized embeddings of the tool definitions and the API documentation:

  • BM25 keyword search catches exact terms such as "submittal", "change order" or "commitment".
  • Semantic embeddings (an ONNX MiniLM model) catch intent when the user's words don't match the endpoint's name.
  • Reciprocal Rank Fusion combines the two ranked lists into one result set.

The search covers more than raw endpoints. Capabilities, workflows, personas and policies are vectorized too, so a request like "close out the open RFIs on this job" can match a multi-step workflow, not just a list of loosely related endpoints.

Ranking also adjusts over time. A Thompson-sampling bandit tunes retrieval per tenant and per persona, driven by a feedback loop from real usage, so ranking can drift toward how each organization and role actually works.

The five-layer architecture

A request passes through five layers on its way from the model client to Procore:

  1. Transport. The main transport is streamable HTTP, served from a Docker container that can be deployed on any remote host. Local hosting also works. Authentication headers go with each request, so one deployment can serve many users at the same time.
  2. Mode. Expanded, hybrid or router, picked from the request headers.
  3. Resolver. Turns intent into specific tools, using the hybrid search, capability definitions and planning described above.
  4. Dispatch. The single path every tool call goes through (covered in the next section).
  5. Generated clients. Typed TypeScript handlers produced from Procore's OpenAPI specification.

The server runs on Node.js and TypeScript. Each client supplies its own Procore client ID, client secret and company ID, and the server's environment variables are configured at deployment.

Code generation keeps the tools in sync with the API

All 2,755 tools are generated from data/combined_OAS.json, Procore's combined OpenAPI specification. A codegen pipeline reads the spec and produces the definitions, handlers and registrations for every tool. When Procore adds endpoints, changes data objects or moves URLs, you run the pipeline again. Validation workflows then check that the regenerated tools still work.

Nobody maintains 2,755 tools by hand, and nobody should try. Generating them is the only way the MCP surface stays aligned with the API it wraps.

One dispatch chokepoint for safety

Every tool call, in every mode, goes through the same dispatch sequence:

  1. AJV schema validation. Reject malformed inputs before they reach Procore.
  2. Rate limiting. Enforce request limits and quotas.
  3. Policy evaluation. Check whether this caller may run this operation (for example, blocking deletes, or allowing writes only in a sandbox).
  4. Idempotency. Stop retries from creating duplicate records.
  5. Handler. Make the actual Procore API call.
  6. Output shaping. Trim and normalize the response so the model gets what it needs instead of raw payloads.
  7. PII redaction. Remove sensitive fields before anything returns to the model.
  8. Audit. Record what was called, by whom and with what result.

Having a single chokepoint matters because a control you have to remember to add to each tool will eventually be missing from one of them. When every call goes through one path, a policy written once applies to all 2,755 tools.

Policies, personas, capabilities and workflows

These four building blocks turn raw API access into something an organization can run. All of them are defined in YAML and can be extended:

  • Capabilities group the tools a task area needs. Working with RFIs, for example, involves several endpoints plus the rules for how they're used together.
  • Workflows are multi-step sequences of tool calls, with the data each step needs.
  • Personas describe who is using the server, so a project engineer and a controller each see a scope that fits their role.
  • Policies enforce what can and can't happen: read-only access, checksum and validation loops before creates or updates, sandbox-only writes, rate limits and quotas.

The walkthrough also covers how policy mechanics support organizations that have to meet standards such as GDPR, HIPAA and SOC 2. The server gives you the enforcement points. Each organization defines what its own standard requires.

Multi-tenant by design

Tenant boundaries come from request headers. One deployment can serve multiple teams and organizations, each in its own partition, so data doesn't cross between them. Async jobs and backend workers keep one heavy user from slowing response times for everyone else. Caching and state are scoped to the session or the organization, depending on what the data is.

Observability, plugins and evaluation

An agent platform you can't observe is one you can't operate. The server exposes an MCP endpoint for streaming requests, plus health and metrics endpoints. Telemetry covers:

  • Uptime and downtime
  • How often each tool is called
  • Which tools fail, and why
  • Which organizations and IP addresses are using the server

The same data shows where to improve next, and it can feed usage-based pricing and availability decisions.

Plugins can register more workflows, personas and resolution logic, and they can apply rules before dispatch. This is where an organization's own requirements, or a platform's API terms, get encoded.

Evaluation and tuning close the loop. Benchmarks check that routing still picks the right tools, and real queries show which capabilities to build next. The bandit-tuned retrieval uses the same feedback.

What this means if you're connecting agents to Procore

If your team is evaluating AI agents on top of Procore, or any construction platform with a large API, these are the questions to ask:

  • How many tools does the model see on each turn? If the answer is "all of them," expect high costs and poor tool selection.
  • Where do credentials live? OAuth tokens and client secrets belong on the server, never in a prompt.
  • Is there one place where policy is enforced? Ask how deletes are blocked and how writes are confined to a sandbox.
  • Can you see what the agent did? Every call should have an audit trail you can review.
  • What happens when the vendor's API changes? Without generated tools, drift is guaranteed.

Practical lessons from the build

  • Make tool counts a non-issue. Generate the full surface, then control exposure with modes. Don't hand-pick a few hundred endpoints and hope they cover every use case.
  • Invest in retrieval before prompts. In router mode, search quality sets the ceiling: the agent can only call what search surfaces. Hybrid keyword and semantic search with rank fusion is a strong default.
  • Encode domain knowledge as data. Workflows and personas in YAML are easier to review and extend than logic buried in prompts.
  • Put every safety control in one path. Validation, policy, idempotency, redaction and audit only work reliably when nothing can go around them.
  • Ship it as a container with health checks. We deploy with Docker because it's quick. Kubernetes or an on-premises service behind a reverse proxy are reasonable alternatives, as long as you keep the health and metrics endpoints.

The same pattern applies beyond Procore. Any large SaaS API you expose over MCP, whether ERP, scheduling or document control, runs into the same context, selection and governance problems.

Where to go next

Frequently asked questions

Why not just expose every Procore endpoint as an MCP tool?

MCP clients send tool definitions to the model on every call, so 2,755 definitions consume a large share of the context window and make tool selection less reliable. It also gives every agent visibility into every write and delete operation. Routing exposes only what the current task needs.

What are MCP meta-tools?

Meta-tools are a small set of tools that let an agent search the full tool catalog, read the details of a candidate tool, and then call it. In our Procore MCP server, router mode exposes about 25 of them, including procore_search_tools, procore_call_tool and procore_plan_dag.

How does the server pick the right Procore tool for a request?

It runs hybrid retrieval over embeddings of tool definitions, API documentation, capabilities and workflows. BM25 keyword search and ONNX MiniLM semantic search are combined with Reciprocal Rank Fusion, and a Thompson-sampling bandit tunes ranking per tenant and persona using feedback from real usage.

How are AI agents kept from doing something unsafe in Procore?

Every tool call passes through a single dispatch path that validates input, applies rate limits, evaluates policy, enforces idempotency, shapes and redacts output, and writes an audit record. Policies can block deletes, require validation before creates, or restrict writes to a sandbox.

Can one Procore MCP server serve multiple companies or teams?

Yes. Tenant boundaries are set through request headers, so multiple teams and organizations can share a deployment while their data, state and caches stay partitioned. Async jobs keep heavy users from slowing everyone else down.

How is the Procore MCP server deployed?

It is a Node.js and TypeScript server that runs as a Docker container over streamable HTTP, with health and metrics endpoints. Each client supplies its Procore client ID, client secret and company ID. It can also run locally, and Kubernetes or an on-premises reverse proxy are workable alternatives.

Next step

Have a problem like this?

Tell us the outcome you need. We'll tell you honestly how we'd approach it, and reply within two business days.

Keep learning