Most multi-tenant SaaS products do not lose control of authorization in one dramatic mistake. They lose it one route at a time. A developer adds an endpoint, copies the check from the endpoint next to it, gets the tenant filter slightly wrong, and nothing fails. Six months later there are three ways of deciding what "admin" means and nobody is sure which one a given screen uses.
When we built Connect, the AI platform that sits on top of Syncify for a scheduling consultancy, we treated that drift as the main thing to design against. Connect has hundreds of routes, a chat agent, an MCP server and a public embed widget, all reaching the same project data. This article walks through how its role-based access control is put together: one authorization function, a small set of hooks, feature packages instead of loose flags, a route manifest that every endpoint must appear in, and tests that make the rules hard to break quietly.
The tenant model: workspaces, memberships and one admin flag
Everything in Connect hangs off a workspace, which is the tenant. A user reaches a workspace through a row in workspace_memberships, and that row carries one boolean: is_admin.
That boolean is the Workspace Admin role. It is deliberately not a row in the roles table. We argued about this, and the reasons are worth repeating because most teams make the opposite call:
- "Admin on everything the workspace has" cannot be stored as grants. If admin were a role with explicit
role_featuresrows, those rows would go stale the moment the operator switched a new feature on for the workspace. The flag resolves to admin level on every entitled feature at read time, so it never drifts. - "Last admin" is a structural rule. Connect refuses to remove or demote the final admin of a workspace. As an editable role, someone could change the role's grants and lock everyone out without touching membership at all.
- It is what single sign-on can seed. On a user's first login, Syncify's workspace-admin flag sets
is_admin. After that, Connect owns it. A later login does not overwrite a demotion made in Connect.
The flag answers "does this person own the workspace?" Roles answer "what does this person do here?" An admin also holds no rows in user_roles: a second role could only grant less than the flag already does.
Features, packages and roles
Under the admin flag sits a conventional RBAC model with three tables:
| Table | Holds |
|---|---|
workspace_features | Which features this tenant is entitled to |
roles / role_features | Named roles, each granting a level (0 = normal, 1 = admin) on some features |
user_roles | Which members hold which roles |
Two details in the schema matter more than they look. The level is a smallint with a check constraint rather than text, because the check is a >= comparison and text compares lexically. And user_roles has a composite foreign key to workspace_memberships, so it is impossible to assign a role to someone who is not a member of that workspace, and the assignment disappears when the membership does.
The feature catalog is code, not a table
The list of features (projects, files, schedules, apps, the AI assistant, users and so on) lives in a TypeScript constant, not in the database. A feature only exists once its sidebar item and routes ship, so a database row would always be a second thing to remember. Feature keys are also referenced in route wiring, where a typo should fail the type checker rather than 403 in production.
Features are sold in packages
Early on, features were toggled one key at a time, which let a workspace hold Schedule Builder without Schedules: not a cheaper product, a broken one.
So features are grouped into packages, and a package can declare requires. A small resolver takes the packages someone asked for, always adds the core package (projects, users and files), and pulls required packages in transitively. It never removes anything. Asking for Apps without Scheduling gets you both, rather than an error.
The direction of each dependency is the point. A builder needs what it builds on, never the reverse. Scheduling without Schedule Builder is a coherent product that real workspaces run, so the edge points one way only. The AI assistant is its own package and is never a requirement of anything: revoking it stops agent turns while manual editing keeps working.
Three tiers, two of them rows
The seeded roles mirror Syncify's tiers so both products describe access the same way:
- Workspace Admin: the
is_adminflag, admin on every entitled feature. - Workspace Member: the default role. Admin level on everything the workspace holds except user management and connectors, which stay at normal level. The policy in one line: members are trusted with the work, not with the tenant's keys or its people.
- Workspace Guest: normal level on Apps, and nothing else. A guest comes in to look at a dashboard built for them, not to browse projects or files.
The single chokepoint: authorize(context, permission)
Every route hook, and every agent-side check, ends up asking one function. In shape it looks like this (written fresh for this article, not copied from the repo):
type Permission = "workspace.read" | "workspace.write" | "system.admin" | "agents.admin"
function authorize(ctx: AuthContext, p: Permission): boolean {
if (p === "agents.admin") return ctx.user.isSuperAdmin && onOperatorAllowlist(ctx.user)
switch (requiredTier[p]) {
case "superAdmin": return ctx.user.isSuperAdmin
case "workspaceAdmin": return ctx.workspaceId !== null && ctx.isWorkspaceAdmin
case "member": return ctx.workspaceId !== null
}
}
A few choices are baked in here:
Named permissions, not booleans. Call sites ask for workspace.write, not isAdmin. Today every permission resolves to a tier, but the signature is the seam. If a third tier arrives, this function changes and the call sites do not.
"Admin of nothing" is denied. An admin-scoped route with no workspace in its path is a wiring mistake, so it fails closed instead of passing.
The executive surface is doubly gated. The operator's agent-governance dashboard requires a super-admin account and an operator-domain allowlist on the email. The email comes from SSO, which Connect trusts, so the domain check is a real boundary rather than a hint.
The rule we wrote in the code comment is the one that matters: route hooks and agent tools call this and nothing else, because two implementations diverge and only one of them gets audited.
The hooks that feed it
authorize only works if the context it receives is honest. A short chain of Fastify hooks builds that context, and each one does exactly one job:
- requireSession resolves the session cookie to a user, or returns 401. It actively clears a stale or tampered cookie, so a bad cookie does not re-fail every request forever.
- sameOrigin is the CSRF defence for cookie-authenticated mutations. Sessions are hand-rolled, so there is no framework to do this. SameSite=Lax blocks cross-site POSTs but not same-site sub-origins, so every non-GET request must carry an Origin (or Referer) that matches the public URL.
- requireMembership reads
{workspaceId}from the path and checks it against the caller's memberships. Not-a-member and no-such-workspace both return 403. A 404 would confirm that the ID exists. The workspace on the context only ever comes from the path, never from a body or header. - requirePermission calls
authorizewith the scope's permission. - requireFeature(feature, level) checks that the workspace is entitled to the feature and that the caller's roles reach the level. A workspace admin resolves to level 1 on everything, so this is never weaker than
workspace.write. - requireProjectAccess applies per-member project scope (below).
Membership runs as a preHandler because path parameters only exist after routing, and it goes through the workspaces module, not its store: one capability never reads another capability's tables, a rule we enforce with tests (covered in the architecture tests article).
Project scope resolves the resource, not the URL
A workspace admin can see every project. A member can be narrowed to a list of projects. The obvious implementation checks {projectId} in the path. That is what we started with, and it does not generalize: by-id routes such as /notes/:noteId or /files/:fileId have no project ID in the URL. The owner is one column away, and a params-only check cannot see a column.
The fix is a ProjectOwners map that sits beside the route manifest. Each path parameter that names a project-owned resource maps to a resolver that returns the owning project IDs:
projectId -> itself
programId -> the projects in that program
noteId -> the project the note is filed under
fileId -> the project(s) that hold the file
...
The hook finds the first matching parameter, looks up the caller's allowed projects (null means unrestricted, so the common case loads nothing), resolves the resource's owners, and denies with the same 403 as non-membership if none of them are in reach. Order in the map matters: the parameters that name the project directly come first, so a route carrying both :projectId and a schedule ID resolves through one indexed read rather than a join.
Where the rules live: route scopes
All of this is wired in one file. Fastify's plugin encapsulation is the safety property: a hook added inside a plugin applies to every route in that plugin and to nothing outside it. So the routes file is a set of scopes, each with its hooks stacked at the top:
| Scope | Credential | Gate |
|---|---|---|
| Public | None, or a token in the path | Handled inside the module |
| Session | Cookie | Valid session only |
| Member | Cookie | Membership, workspace.read, project scope, then per-group feature gates |
| Apps | Cookie or personal access token | Same as member, plus PAT pinned to its workspace |
| Workspace admin | Cookie | workspace.write |
| Embed | Embed bearer token | Token scope only |
| System / executive | Cookie | system.admin or agents.admin |
Inside the member scope, a helper wraps each group of routes in its own nested plugin with a requireFeature hook. That helper exists to make one mistake impossible: adding a feature hook on the parent scope and accidentally gating everything in it. The rule in the code is blunt. Never add an authorization hook at the root, because that would cover the public routes too.
The manifest: every route declares an auth level
At the top of the routes file is MANIFEST, a plain array of [method, path, level] tuples. It holds several hundred rows. The level is one of unauthenticated, session, member, workspaceAdmin, superAdmin, embed, execAdmin, or a feature gate written as feature:level, such as projects:1.
The manifest is not documentation. test/routes.test.ts builds the real app and asserts several things about it:
- The registered route set equals the manifest. Not a subset. A slice cannot grow a route that never reaches review, and a deleted route cannot leave a stale row.
- Every feature-gated route registers at the level the manifest claims. The feature helper tags each route with the gate it actually registered under, and the test compares tags with manifest levels. This test exists because we once found a handler registered in a lower-level group than its manifest row claimed, and the earlier tests had no way to notice.
- Every path parameter is classified. Each
:parammust appear inProjectOwnersor in an explicit, commented "not project-owned" list. Both lists are checked in the other direction too, so a removed route leaves no stale entry. - Every non-public route returns 401 without a session, and a plain member with no roles gets 403 from every admin and feature-gated route.
Unauthenticated rows carry a comment naming the credential that replaces the session, such as a signed transient cookie on the Procore OAuth callback or an unguessable token in a webhook path.
The third door: agents and MCP
Feature gates are route hooks, and an AI product has more doors than routes. Connect's settled decisions record this as a sequence.
The chat agent's tools were the second door. They call the same capability modules, so without their own gate a member without schedule access could ask the assistant for the portfolio instead. So the tool scope carries a required features map, and the tool list is filtered before the model sees it. A tool the model never sees cannot be talked into running. Level 0 is enough, so the agent is exactly as capable as the caller's own UI.
MCP was the third door, and it was gated at the same chokepoint: the MCP module's workspace and project checks now take the features each tool reads. A missing feature answers "workspace not found", the same as non-membership, so an MCP client never learns which features a workspace has. More on that surface in building an enterprise MCP server with OAuth.
The phrase "third door" also shows up in a subtler decision. Role authoring requires the admin flag, not users:1. At users:1, a delegated user manager could author a role granting admin on everything and assign it to themselves: the same self-escalation that member creation and promotion already forbid, through a third door. So users:1 keeps the delegated half (adding members, assigning roles an admin authored), and defining what access exists stays with the owner. The test we use for every such call: "if this feature were revoked, should this action stop working?" Yes means a feature level. No means is_admin.
The access_changes ledger
Every RBAC mutation writes a row to access_changes: workspace, subject, who changed it and when. The table is append-only, and the latest row per subject wins for attribution. Setting features, creating, editing or deleting roles, assigning roles and narrowing project scope all take a trailing actor ID, and the admin screens show "updated by" from the ledger. A deleted actor leaves the row with a null user rather than erasing the history.
It is a small table that answers the first question after any access incident: who changed this, and when?
Why this matters if you're building something similar
- Write one authorization function and make it the only one. Name the permissions. The signature is the seam you will need later.
- Deny with the same status everywhere. Non-member, wrong project and missing feature should all look identical from outside, so none of them can be used to probe.
- Scope by the resource, not the URL. If a route takes an ID, something must resolve who owns that ID. Make it a map and test that every ID parameter is in it.
- Keep ownership separate from capability. An admin flag on membership and roles for everything else is simpler than an admin role that has to be kept in sync.
- Sell coherent bundles. Packages with one-directional
requiresprevent configurations that are cheaper on paper and broken in practice. - Make the manifest executable. A list of routes and auth levels is only useful if a test fails when the code disagrees with it.
- Count your doors. Routes, agent tools, MCP and embeds should all reach the same chokepoint. Every new surface is a new door.
Where to go next
- The series overview: How we built Connect, an enterprise AI platform
- How other credentials reach the same rules: Headless auth: personal access tokens and embed tokens
- The tests that keep this from eroding: Architecture tests that enforce codebase boundaries
- The agent-side counterpart: Agent tool manifest and permissions
- The governance view for leadership: Governing AI agents in construction
Planning a multi-tenant platform of your own? Plan your build with us.
Frequently asked questions
Why return 403 instead of 404 for a workspace the user cannot access?
Returning 404 for a missing workspace and 403 for one you are not in tells an attacker which IDs exist. Returning the same 403 for both, and for out-of-scope projects and missing features, gives nothing away.
Why is Workspace Admin a flag and not a role?
Admin means admin on every feature the workspace has, including ones added later. Stored grants would go stale, and an editable role could be changed to lock everyone out. A boolean on membership resolves at read time and supports a last-admin rule.
What is a route manifest?
A list of every API route with its required auth level, kept in one file. A test builds the real app and fails if the registered routes differ from the manifest or register under a different feature level.
How do feature packages differ from feature flags?
Packages group related features and declare one-directional requirements, so a tenant cannot be entitled to a builder without the data it builds on. A resolver always adds core features and pulls requirements in transitively.
How do AI agents fit into RBAC?
Agent tools and MCP endpoints reach the same capability modules as routes. Connect filters the tool list by the caller's feature access before the model sees it, and the MCP server applies the same membership and feature checks.
Next step
Need something built around how your team works?
Describe the users, the workflow, and the systems it touches. We'll tell you whether a custom application makes sense and how we'd build it.
Prefer email? charley@buildflows.ai
Get the next guide in your inbox
Field Notes: practical guides and new walkthroughs, about once a month.
Field Notes
Practical guides and new walkthroughs on construction data and automation, roughly monthly.
Keep learning
AI agents & MCP · October 9, 2026
How We Built Connect: Architecture of an Enterprise AI Platform
The pillar of our Connect architecture series: a layer-by-layer map of an enterprise agentic AI platform, the three-process shape it runs as, the design principles that kept recurring, and links to every deep-dive article.
Custom applications · October 9, 2026
Headless Auth: Personal Access Tokens, Embed Tokens and One Door per Credential
How Connect handles browser sessions, CLI personal access tokens, cross-site embed tokens and MCP OAuth, with each credential confined to one route scope and resolved to the same authorization rules.
Playbooks · October 9, 2026
Architecture Tests: Enforcing Codebase Boundaries So Small Teams and AI Agents Can Move Fast
The tests that govern Connect's structure: import layering, a route manifest checked for equality, catalog parity, agent tool rules, design tokens, migrations and a rehearsed deploy script.