Build Flows

Integrations · October 9, 2026 · 11 min read

Building a Connector Framework for Procore and OAuth Integrations

How Connect's connector framework is built: compiled-in definitions shared by API and worker, credentials sealed and bound to their connection, race-safe token rotation, leased sync runs and a clean split between source and manual project links.

By Charley Forey, founder of Build Flows

An AI platform for construction is only as useful as the data it can reach. For Connect, the AI platform we built on top of Syncify, that meant Procore projects, RFIs and submittals, Autodesk and Bluebeam documents, Google Drive and OneDrive folders, weather for each jobsite, and Syncify's own schedules. Each source has its own auth model, rate limits and ideas about what a "project" is.

We have written before about the Procore API itself and about routing an agent across thousands of Procore endpoints. This article is about something else: the connector framework inside a multi-tenant product. That means how connectors are declared, how credentials are sealed, how sync runs are leased, how synced data maps onto projects, and the decisions we would make again.

One registry, used by the API and the worker

Every connector is a compiled-in TypeScript definition. A single function builds the registry, and both the API process and the background worker call it.

That removes a whole class of bug. A separate API catalogue and worker dispatch table drift apart, and you get a connection an admin can create that nothing can ever sync. With one registry, a connector can't be offered unless it can also be run.

We also decided against dynamic loading: no third-party connectors, no plugin directory, no sandbox. Node makes it easy to drift into, so we wrote the decision down.

A definition declares:

  • Identity and copy: name, display name, description, and mandatory attribution (public data sources have attribution terms).
  • Auth mode, as a discriminated union: none, credentials (username/password or API key), oauth, or syncify_key (a key that arrives through Connect's own SSO round trip and always rotates).
  • Scopes: what a connection may import, such as a Procore company and its projects, and how many.
  • Tables: the shape of the records it produces.
  • Setup guide: the provider hosts Connect will call and why, the OAuth scopes a customer's app must grant, and the steps to register one. This sits next to the code that calls those hosts, so a customer's security review reads facts that ship with the build.
  • Flags for whether it produces files and whether it honours an incremental since cursor.

The auth mode is a type, not a runtime check. An OAuth definition must supply authorizeUrl and exchangeCode, a credentials definition must declare authentication fields, and an OAuth definition must declare none, because in OAuth the user types their secret into the provider, never into Connect. The compiler enforces the shape, and the registry validates it again when the process starts.

JSON table schemas, validated at boot

Table shapes live in JSON files, one per connector. At startup the loader:

  1. Parses every schema and rejects duplicate tables, duplicate columns, and key columns that don't exist.
  2. Adds connection_id and synced_at to every table. Multi-tenancy comes from the platform, so no connector author has to remember it.
  3. Compares each definition's declared tables with its schema file and refuses to start if they differ.

Startup failures are the right severity here, because these files ship with the build and user input can't produce them. A wrong schema should stop the deploy, not render a broken form for a customer.

The frontend's connector picker renders from a serialized form of the same definitions. The executable source object never leaves the process, so there's no hand-maintained copy of the schema in the browser.

For Procore we store about ten curated tables (companies, projects, RFIs, submittals, direct costs, meetings, prime contracts, change events, punch items, observations). Each has lean typed columns plus a data column holding the whole provider object. App builders get a clean schema to query, and agents still have every field.

Credentials: a separate table, sealed with AES-256-GCM

Connections and their secrets live in different tables. A connections row can be listed, serialized and sent to the browser freely, because no read path joins to connection_secrets. A credential can't ride along in a response by accident when the query that would include it doesn't exist.

Each secret row stores four values: key_version, nonce, ciphertext and auth_tag. The token payload is encrypted with AES-256-GCM, using a random 12-byte nonce each time it's sealed. The important part is the additional authenticated data: workspace id, connection id and connector name, joined with NUL separators. The ciphertext is cryptographically bound to the connection it belongs to. Copy one connection's sealed blob into another connection's row and it fails to decrypt. It doesn't quietly hand one tenant's Procore token to another tenant.

The key_version column exists so the key can be rotated. Today there's one version, and a credential sealed under a version this build doesn't know fails with a specific error instead of a generic decryption failure. That gives a key rotation a clear failure mode instead of a confusing one.

In illustrative form:

sealed = aesGcmEncrypt(key, randomNonce(12), JSON(tokens),
                       aad = [workspaceId, connectionId, connector].join("\0"))
store(connectionId, { keyVersion: 1, nonce, ciphertext, authTag })

Refresh-token rotation without races

Several providers rotate refresh tokens, and Syncify adds reuse detection: replaying a retired refresh token revokes the whole session. The two failure cases that matter are two workers refreshing at the same moment, and a refresh call that times out after the provider already rotated.

The rules we follow:

  • Only use a token that was successfully written. After a refresh, the new sealed secret is written with a compare-and-swap on the previous ciphertext. If the swap loses, another worker rotated first, so we re-read and adopt their token.
  • Never retry a refresh that may have succeeded. If the refresh call fails, we re-read the stored secret before doing anything else. If it changed, someone else rotated and we use theirs. Retrying would replay a consumed token.
  • Carry the app secret across rotations. For bring-your-own OAuth apps, the customer's client secret is sealed in the same payload as the tokens and preserved on every refresh.

A refresh failure that means the user must sign in again moves the connection to reauthorization_required and notifies the workspace. That's different from disconnected (added in migration 002), which is for a connection someone removed whose ingested data should stay: the row survives with no key, and everything hanging off it is kept. Both are different again from deleting, which means a purge is in progress.

Procore OAuth: shared app or your own

Procore uses the authorization code grant with refresh-token rotation. A workspace can connect in two ways:

  • Through Connect's shared Procore app. The deployment holds one client id and secret, and both must be configured together or not at all, checked at boot. If they aren't configured, Procore isn't offered.
  • With its own app. Some enterprises require their own registered app for security review. The customer's client id goes in the connection's configuration in the clear. The client secret is sealed into connection_secrets next to the tokens. There's no separate table and no plaintext secret.

Either way, the auth interfaces take the app credentials as an argument that the connections module resolves per call, so exchange, refresh and revoke always use the same app.

All OAuth connectors share one fixed callback URL, with no workspace id in the path, so a provider needs only one redirect URI registered. Here is the round trip:

  1. An admin creates the connection. The row is created in status authorizing.
  2. The server sets a short-lived signed cookie binding workspace, connection, user and a random state value, then redirects to the provider.
  3. The provider redirects back to the shared callback with a code and the state.
  4. The callback trusts only the signed cookie and the state match, never the query string, to decide which connection the code belongs to.
  5. The code is exchanged, the tokens are sealed, scopes are discovered, and the connection becomes active.

The signed cookie is what ties the callback to the browser that started the flow: a code that arrives without the matching cookie and state is rejected before any exchange happens. Connect's own sign-in and its MCP server add S256 PKCE on top of the same idea.

A connection can also target Procore's sandbox instead of production on a per-connection basis. Each environment gets its own HTTP client, because the pacing clock is per host.

Live, per-project reads for agents

Synced records are the cheap, fast path, and agents read them first. Sometimes an agent needs fresher data, or a Procore endpoint outside the curated tables. For that there's a read-only tool that issues one GET against Procore's REST API on behalf of a Connect project. The server pins the path to that project's Procore project id or its company, and refuses paths for other projects or companies. It injects the company header and bearer token, and refreshes the token if it's stale. The agent never sees a credential and can't write. See the agent tool manifest for how tools like this are gated.

Procore connections also register company-level webhooks for the resources we sync. Each delivery is authenticated per connection and does only one thing: it triggers an incremental sync for that connection. Several deliveries in a row collapse into one sync.

Leased sync runs

A sync is a row in connection_runs, and most of the worker's correctness lives in Postgres indexes, not in application code:

  • One active run per connection is a partial unique index on connection_id for runs that are queued or running. Two workers and three browser tabs racing to start a sync produce exactly one run, because the insert is ON CONFLICT DO NOTHING.
  • Claiming is leasing. A worker claims a run by setting lease_owner and lease_expires_at, and renews the lease while it works. A running row whose lease has expired can be claimed again. That is crash recovery. There's no separate reaper process.
  • Losing the lease means stop. If a renewal finds another owner, the worker stops writing at once. Two workers writing one connection's records is exactly the corruption the lease exists to prevent.
  • Retries reuse queued_at as "not before", so there's no extra scheduling column.

The leased background jobs article covers the claim query in depth.

Inside a run, the worker walks the connector's declared tables, not whatever tables the connector returned. A connector can't skip a table and leave stale data behind. Records land in connector_records, keyed by connection, table name and external id, with the provider row stored whole as jsonb. Change detection uses a content_hash computed in Postgres, because jsonb normalizes key order and JSON.stringify doesn't: the same row with its fields reordered would otherwise look changed on every sync. A full sync stamps one synced_at per run and soft-deletes rows it didn't see. An incremental sync passes the since cursor and only upserts.

One failing document doesn't stop the other nine hundred. File failures are counted and the run reports partial instead of throwing.

Mapping data onto projects: SOURCE vs MANUAL links

The link table connection_project_links joins a connection to a Connect project. Early on it had one rule: one link per project. That was right for sources and wrong for everything else.

Migration 028 split links into two kinds:

KindExampleRule
SOURCEA Procore or Syncify project that a sync projects into ConnectAt most one per project, enforced by a partial unique index WHERE NOT manual
MANUALA weather forecast, a Drive or OneDrive folderAny number, each keyed by the Connect project and carrying its own configuration

Before the split, assigning weather to a Procore-sourced project hit the unique constraint and showed up as a generic internal error. Afterwards, enrichment links stack freely beside the one real source. The ownership rules (who can overwrite project details, which link a refresh trusts) count only sources.

Folder-linked connectors. Google Drive, OneDrive, Autodesk and Bluebeam are bring-your-own OAuth: the customer registers their own app with the provider, and the setup guide lists the redirect URI, scopes and hosts. Drive and OneDrive don't discover projects at all. You connect once, browse folders, and link a folder to a Connect project as a manual link. The sync then runs one extract pass per linked folder and tags each streamed file with its project, so it lands in that project's Files.

A Connect-managed weather key. AccuWeather runs on a key Connect holds. The workspace connects with no credentials and no secret row, and the connection activates immediately. Weather is assigned per project as a manual link carrying the location, and the sync extracts once per link. A project with no location returns location_required so the UI asks for one. Open-Meteo is offered alongside it as a keyless option with a longer forecast horizon.

Syncify syncs itself daily (migration 059). Syncify connections used to sync only when someone clicked the button. Activation now sets a daily schedule, and the migration moved existing active connections from off to daily. It left any schedule someone had chosen deliberately alone, and seeded next_run_at to now so each connection got one catch-up run instead of a backlog.

The Suggested Connectors board

We didn't want to guess which connector to build next. Migration 014 added a platform-wide request board: any signed-in user can file a request and upvote others, and the submitter's vote is added automatically. Requests display anonymously and sort by votes. The submitter is stored for dedupe and moderation but never returned by the API, so the board measures demand without turning into a leaderboard of people. An executive admin moves each request through requested, considering, planned, in_progress, live_soon, live or declined.

Tests run against stubs, always

The connector test suites never call a real provider, and there's no "live mode" switch for anyone to flip in CI. Connector clients are built from base URLs, not hard-wired hosts, so a test can point one at a local stub. That makes the suites deterministic. A vendor outage or rate limit can't turn the build red, and no test needs a real customer credential. The cost is that a stub only knows the provider shapes we taught it. When a provider changes its API, a passing test doesn't prove the connector still works against the real thing.

Why this matters if you're building something similar

  • Use one registry for "can create" and "can run". It rules out connections that can never sync.
  • Keep secrets in their own table and bind the ciphertext to its owner with AEAD additional data. Version the key from day one.
  • Treat refresh as a compare-and-swap. Never retry a refresh that might have succeeded. Re-read and adopt.
  • Let the database enforce concurrency. A partial unique index plus a lease is simpler and more reliable than coordinating in application code.
  • Separate sources from enrichments. "One per project" was right for sources only, and modelling both as one thing produced a confusing error.
  • Put setup facts next to the code that calls the hosts, so security reviews read the truth.

Where to go next

Need connectors into your own systems? Plan your build.

Frequently asked questions

How should a SaaS product store customers' OAuth tokens?

Keep them in a separate table that no ordinary read path joins, encrypt them with an AEAD cipher such as AES-256-GCM, bind the ciphertext to its tenant and connection with additional authenticated data, and record a key version so the key can be rotated.

Can a customer use their own Procore app instead of a shared one?

Yes. In Connect a workspace can connect through a shared Procore app or supply its own client id and secret. The client id is stored in configuration and the secret is sealed alongside the tokens and preserved through every refresh.

How do you avoid refresh-token reuse errors with concurrent workers?

Write the rotated credential with a compare-and-swap on the previous ciphertext and use the new tokens only if the write succeeded. If a refresh call fails, re-read the stored credential and adopt any newer token instead of retrying.

How do AI agents read live Procore data safely?

Connect exposes a read-only tool that issues a single GET on behalf of one Connect project. The server pins the path to that project's Procore project or company, injects the credentials and refuses other paths, so the agent never sees a token and cannot write.

What is the difference between a source link and a manual link?

A source link is the one provider project a sync projects into a Connect project, limited to one per project. Manual links are enrichments such as a weather forecast or a linked Drive folder, and a project can have any number of them.

Next step

Need this connection in your environment?

Scope one integration: the records, direction, timing, and business rules behind the connection.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning